Pith. sign in

REVIEW 4 major objections 5 minor 9 references

Easy and intermediate CTF challenges in cryptography, web, and binary exploitation are now reliably solved end-to-end by frontier AI agents, so undeclared scoreboards no longer measure human skill.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 02:26 UTC pith:PDKUUV43

load-bearing objection The purpose-first CTF framework is genuinely useful, but the headline claim that easy/intermediate crypto, web, and pwn are 'reliably automated' outruns the paper's own evidence. the 4 major comments →

arxiv 2607.25425 v1 pith:PDKUUV43 submitted 2026-07-28 cs.AI cs.CRcs.CY

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play

classification cs.AI cs.CRcs.CY
keywords Capture the Flaglarge language modelsagentic systemscybersecurity educationassessment integritycapability boundarycompetition designfair play
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that the disruption of Capture the Flag competitions by large language models is real, category-specific, and already reshaping what scoreboards mean: frontier agents now reliably automate easy and intermediate cryptography, web exploitation, and binary exploitation challenges, while narrow sub-categories (novel mathematics, bespoke logic flaws, physical or social components) still resist. The authors argue that the community's polarized debate over whether AI is 'cheating' is downstream of an undeclared prior question: what a given competition is for. They contribute a map of the current human-machine capability boundary, qualitative evidence from players and organisers, and a four-component safeguard framework whose combination is tied to a competition's declared purpose. A sympathetic reader would care because if the empirical claim is right, CTF rankings in these categories no longer certify human skill, and the same confusion now threatens exams, certifications, and hiring in cybersecurity.

Core claim

The paper's central discovery is that, as of mid-2026, frontier agentic LLMs reliably solve easy and intermediate challenges in cryptography, web exploitation, and binary exploitation—often at speeds no human can match—while a narrow set of sub-categories (novel mathematics, bespoke logic flaws, race conditions, top-difficulty kernel work, physical or social components) still resists. Observed live-competition solves agree with published benchmarks and show that the agentic transition mattered more than model scale; each resistant sub-category erodes once write-ups enter training data. The authors conclude that without a declared purpose, CTF scoreboards quietly collapse into a single outcom

What carries the argument

The central object is a purpose-first safeguard framework rather than a single technique. It consists of four components—tiered competition divisions, LLM-resistant challenge design, telemetry used investigatively rather than adjudicatively, and a draft community code of conduct built on post-competition disclosure—plus a decision tree that maps a declared competition purpose (learning-first, ranking-and-qualification, or hybrid) to a weighting of those components. The conceptual hinge is the distinction between AI as a cognitive prosthesis (which speeds up a player who still directs the work) and AI as cognitive offloading (which replaces the reasoning and leaves the player as a validating

Load-bearing premise

The load-bearing premise is that the observed solves were genuinely AI-driven and that the handful of events studied represent modern CTFs generally; if the AI attribution is inflated or the sample unrepresentative, the claim that easy and intermediate challenges are 'reliably automated' collapses, even though the conceptual framework could survive.

What would settle it

Take a random sample of 50 easy-to-intermediate cryptography, web, and pwn challenges from public 2025 CTFs, run a single frontier agent with a fixed token budget and no human intervention, and record solve rate and wall-clock time; if the agent solves far less than the 'reliably automated' share reported here—or requires human debugging to reach a flag—the empirical boundary is falsified. The same test could be repeated quarterly to track the claimed doubling of reliable task length.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the capability map is right, standard easy and intermediate crypto, web, and pwn challenges are no longer valid evidence of human ability, so scoreboards in those categories mislead everyone who reads them as skill rankings.
  • The resistant sub-categories are a moving target: any design that works this season erodes once it is written up and absorbed into training data, so LLM-resistant design is a recurring cost, not a one-time fix.
  • The agentic transition matters more than model generation: harnesses that let models drive tools and recover from errors raise solve rates across all categories, with the largest gains in web and the smallest in genuinely novel mathematics.
  • A competition's AI policy is only coherent after its purpose is declared; the same safeguard used in a learning-first event (permissive AI) inverts in a ranking event (restricted, enforced, telemetry-backed).
  • The argument extends beyond CTFs to any setting where a demonstrated result is treated as evidence of ability—coursework, practical exams, certification, portfolio hiring—and the same declare-purpose-then-choose-safeguards logic transfers directly.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the paper's capability map implies a cheap, continuously updated probe for organisers—serve a small held-out set of known-resistant challenge types each season and measure frontier-agent solve rate, turning the boundary into a live indicator rather than a yearly case study.
  • Editorial inference: the purpose-first logic transfers to certification and hiring, but the paper leaves open how an organisation audits whether its own declared purpose is honest; that is itself a governance problem the framework does not solve.
  • Editorial inference: if the step-change is driven mainly by context-window growth, telemetry heuristics keyed to implausibly fast solves will only get noisier; process-based attestation (write-ups, interviews, disclosure norms) is likely to outlast timing-based detection.
  • Editorial inference: the 'undisclosed AI use' framing suggests a testable social intervention—run one season with mandatory post-competition disclosure for all prize-eligible teams and one without, then compare behaviour and morale; the paper does not run this experiment.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports a mixed-methods study of the impact of large language models on Capture the Flag (CTF) competitions. It synthesizes published benchmarks (Cybench, NYU CTF Bench, InterCode) plus a UK AI Security Institute evaluation, presents observational case studies of live events (DiceCTF 2026, Insomni’hack, national qualifiers), analyzes public community discourse, and reports three semi-structured interviews. The central empirical claim is that easy and intermediate challenges in cryptography, web exploitation, and binary exploitation are now reliably automated by frontier agents, while narrower sub-categories (novel mathematical cryptography, bespoke business-logic flaws, top-difficulty kernel exploitation) remain human-relevant. The paper argues that community disagreement about whether AI use is cheating is downstream of an undeclared question about a competition’s purpose, and it proposes a four-component safeguard framework (tiered divisions, LLM-resistant challenge design, telemetry-based detection, and a community code of conduct) plus a decision tree linking safeguard choices to declared purpose.

Significance. If the capability-boundary claim is sound, the paper identifies a real and timely problem: standard CTF scoreboards in major categories may no longer measure human cybersecurity skill. The conceptual contribution—that a competition’s purpose must be declared before AI policy can be coherent—is valuable and transferable to certification, recruitment, and academic assessment. The four-component framework is pragmatic and appropriately non-prescriptive, and the authors are unusually explicit about limitations, ethics approval, and the dual role of generative AI in the study. However, the paper’s empirical foundation is weaker than its headline claim requires: the case studies are observational, the attribution of solves to autonomous agents versus human-assisted use is not systematically established, and one load-bearing independent evaluation is not verifiable from the reference given. These issues affect the central claim but are addressable through careful reframing and additional evidence, so the paper merits major revision rather than rejection.

major comments (4)
  1. [§4 and Table 1] The central claim that easy and intermediate challenges in crypto, web, and pwn are “reliably automated” is not supported by the evidence as presented. The live-competition examples mix autonomous agents with LLM-assisted human teams: §4.1 says “LLM-assisted teams first-blooded” and “teams … openly credited their LLM agent,” §4.2 says the top team “solved 27 of 30 challenges using AI,” and §4.3 says “LLMs solved almost the entire pwn set”—but no protocol distinguishes solves caused by an autonomous agent from solves where a human drove the reasoning and the model wrote code. The only fully autonomous evidence is the authors’ own platform experiment and a single auto-submit (§2). Without an attribution method, baseline solve rates, or a sampling protocol, “reliably automated” overgeneralizes. Please rephrase Table 1 and the abstract to distinguish “automated” from “AI-assisted,” or add a
  2. [§4.4, Ref [7]] The UK AI Security Institute evaluation is a load-bearing independent datapoint (“closes off the objection that the disruption is overstated”), but the reference is unverifiable: it gives only an institution and URL, with no report title, date, or identifier. The claimed result that “the length of cyber task a frontier model can complete at high reliability has been doubling every few months” is stated with no numeric support. If the report is public, cite it properly; if not, this should be labeled as a personal communication or preprint, and the paper should not rest its central corroboration on an inaccessible source.
  3. [§4.2, Insomni’hack] The inference from the three challenges that resisted the top team to “the remaining frontier is moving toward physicality, novelty, human context and messy interaction” is overinterpretation of n=3. A physically distributed 8 GB challenge, a console video game, and an in-person social-engineering challenge are not a representative sample of hard modern CTF challenges. This is an ancillary claim, but it is used to support the capability-boundary map in Table 1. Please either label this explicitly as an illustrative anecdote or provide additional evidence that these three properties characterize the surviving frontier.
  4. [§5.2 vs. §3 Limitations] The paper says in §5.2 that the three interviewees’ convergence on the purpose question is “the empirical counterpart of this paper’s central claim,” but §3 correctly notes that one participant had been given an outline of the study’s framing before recording. That participant’s convergence is therefore not independent corroboration. The main text should not present the convergence as independent evidence; at minimum, it should explicitly remind the reader of the pre-briefing at the point of the claim, not only in the method section.
minor comments (5)
  1. [§7.1] The text refers to “The Kalmar side-board in Section 7.2,” but KalmarCTF is described in §7.1. Please correct the cross-reference.
  2. [§7.2] “RitzSec and Midnight Sun slop-resistant challenges” appears to be a typo; should read “LLM-resistant challenges.”
  3. [§4.4 / References] Ref [7] needs a full bibliographic entry: report title, authors, date, and stable URL. As written, it is indistinguishable from the institution’s homepage.
  4. [§2] The “timeline of five overlapping phases” in Figure 1 and the accompanying narrative cite no sources for event dates or the “$8,000 in model credits” report. Please add citations or mark these as author observations.
  5. [§6] The illustrative walkthroughs (shellcode constraint, XOR-encoded password, PNG-with-7z, maritime OSINT) are useful but are presented as anecdote. Consider labeling them more explicitly as selected examples rather than systematic data, or adding a short statement of how they were chosen.

Circularity Check

1 steps flagged

No significant circularity: central capability-boundary claim rests on external benchmarks and live-competition observation; only a disclosed interview-framing caveat.

specific steps
  1. other [Section 3, Limitations paragraph]
    "one participant had been given an outline of the study’s framing before recording, so that participant’s convergence with our central question is best read as informed elaboration rather than as independent arrival at the same view."

    The interview strand is presented as 'the empirical counterpart of this paper’s central claim' (Section 5.2), but for one of three participants the convergence is explicitly acknowledged to be informed by the study's own framing. That portion of the qualitative evidence is therefore not independent corroboration; it partly reflects the paper's own question. The circularity is minor and disclosed: the paper flags it in the Limitations paragraph, and the central claim also rests on external benchmarks and community observation, so it is not load-bearing.

full rationale

The capability-boundary claim (RQ1, Table 1) is not circular. 'Reliably automated' is operationalized as observed solves by frontier agents in live competition and is corroborated by external benchmarks (Cybench, NYU CTF Bench, InterCode, and the UK AISI evaluation). No parameter is fitted and then renamed as a prediction; no uniqueness theorem is imported from the authors' prior work; and the reference list contains no self-citations. The benchmark strand is independent and the observational case studies are presented as evidence rather than derived from the framework. The only circular element is explicitly disclosed: one of three interviewees was given the study's framing before recording, so that participant's convergence with the 'purpose question' is informed elaboration rather than independent corroboration. The paper states this itself and does not rest the central claim on that interview alone. The skeptic's concern that live-event evidence may conflate LLM-assisted human solves with autonomous agent solves is an evidentiary/validity critique, not a circularity, and under the hard rules does not raise the circularity score. Verdict: score 1.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 0 invented entities

The paper introduces no new physical or mathematical entities. The axioms it relies on are domain assumptions about the representativeness and accuracy of external benchmarks and reports, the interpretative validity of philosophical frameworks, and the untested effectiveness of its own proposed safeguards. These are reasonable but not established, which contributes to the medium correctness risk.

axioms (6)
  • domain assumption The cited CTF benchmarks (Cybench, NYU CTF Bench, InterCode, CTFAgent) accurately represent frontier model and agent capabilities in live competition.
    Section 4.4 uses these benchmarks to corroborate the observational boundary; if the benchmarks are stale or not representative of live conditions, the boundary map weakens.
  • domain assumption The UK AI Security Institute's 2026 cyber capability evaluation exists as described and its reported task-length-doubling trend is accurate.
    Section 4.4 treats this as an independent state-backed measurement; the report is cited but not independently verifiable in this preprint.
  • domain assumption Solves observed in live events are attributable to LLM/agent use rather than conventional human tooling or misattribution.
    Section 4's case studies rely on scoreboard timing and team claims; no independent verification or random audit is provided.
  • domain assumption The cognitive-prosthesis versus cognitive-outsourcing distinction (Clark & Chalmers; Risko & Gilbert) is a valid framework for interpreting the impact of AI on skill development.
    Section 5.2 uses this distinction to interpret interviews and to argue why LLM agents are different from earlier tools; it is a philosophical framework, not an empirically established theory.
  • domain assumption The Kruger-Dunning illusion-of-competence effect applies to AI-assisted beginners in the way the paper assumes.
    Section 6.1 invokes it to predict pedagogical harm; the paper presents no measurement of this effect in the CTF setting.
  • ad hoc to paper The proposed four-component safeguard framework will have the intended effects when implemented.
    Section 7 proposes the framework and Section 8 admits it has not been evaluated in a live competition; its effectiveness is assumed for the prescriptive claims.

pith-pipeline@v1.3.0-alltime-deepseek · 16023 in / 10081 out tokens · 108692 ms · 2026-08-01T02:26:53.071056+00:00 · methodology

0 comments
read the original abstract

Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web exploitation, and binary exploitation. Large language models (LLMs) can now solve a growing share of challenges with minimal human input, raising urgent questions about fairness, the validity of rankings, and whether participation still delivers the learning that justifies the effort. This paper reports a mixed-methods study of LLM impact on modern CTFs, combining a synthesis of published benchmarks, including a recent government evaluation, case studies of live competition across three challenge categories, structured observation of the public channels where the community debates AI use, and semi-structured interviews with experienced players and organisers. We map the current human-machine capability boundary by category, showing that easy and intermediate challenges in cryptography, web, and binary exploitation are now reliably automated while narrower sub-categories continue to resist. We find that community disagreement about whether AI should be permitted is downstream of an undeclared prior question: what a competition is for. Against this backdrop we contribute a four-component safeguard framework, combining tiered competition divisions, LLM-resistant challenge design, telemetry used investigatively, and a draft community code of conduct, together with a decision tool that ties the combination of safeguards to a competition's declared purpose. The argument reaches beyond CTFs to any setting in cybersecurity where a demonstrated result is taken as evidence of an underlying ability.

Figures

Figures reproduced from arXiv: 2607.25425 by Guo Gen Ang, Harmony Bouabid, Michael Macaulay, Sasha Shaw.

Figure 1
Figure 1. Figure 1: Timeline of the disruption, plotting the model and tooling releases that drove it (above each axis) [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: An indicative map of existing approaches, locating events by their AI policy on the horizontal axis [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Decision tool for competition organisers. The root question, what is this competition primarily for, [PITH_FULL_IMAGE:figures/full_fig_p018_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

9 extracted references · 3 linked inside Pith

  1. [1]

    Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell

    Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. InProceedings of the 2021 ACM Conference on Fairness, Ac- countability, and Transparency (FAccT ’21). Association for Computing Machinery, New York, NY, USA, 610–623. doi:10.1145/3442188.3445922

  2. [2]

    Andy Clark and David Chalmers. 1998. The Extended Mind.Analysis58, 1 (1998), 7–19. doi:10.1093/analys/58.1.7

  3. [3]

    Zhuoran Ji, Daoyuan Wu, Wenyi Jiang, Pingchuan Ma, Zongjie Li, and Shuai Wang. 2025. Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges. InProceedings of the 2025 ACM SIGSAC Conference on 20 Macaulay et al. Computer and Communications Security (CCS ’25). Association for Computing Machinery, New York, NY, USA, 603–617. d...

  4. [4]

    Justin Kruger and David Dunning. 1999. Unskilled and Unaware of It: How Difficulties in Recognizing One’s Own Incompetence Lead to Inflated Self-Assessments.Journal of Personality and Social Psychology77, 6 (1999), 1121–1134. doi:10.1037/0022-3514.77.6.1121

  5. [5]

    Risko and Sam J

    Evan F. Risko and Sam J. Gilbert. 2016. Cognitive Offloading.Trends in Cognitive Sciences20, 9 (2016), 676–688. doi:10.1016/j.tics.2016.07.002

  6. [6]

    Minghao Shao, Sofija Jancheska, Meet Udeshi, Brendan Dolan-Gavitt, Haoran Xi, Kimberly Milner, Boyuan Chen, Max Yin, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami, Ramesh Karri, and Muhammad Shafique. 2024. NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security. InAdvances in Neural Information Proces...

  7. [7]

    2026.Advanced AI Evaluations: Cyber Capabilities Update

    UK AI Security Institute. 2026.Advanced AI Evaluations: Cyber Capabilities Update. Technical Report. UK AI Security Institute. https://www.aisi.gov.uk/

  8. [8]

    John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao. 2023. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. InAdvances in Neural Information Processing Systems 36 (NeurIPS 2023, Datasets and Benchmarks Track). Curran Associates, Inc., Red Hook, NY, USA. doi:10.48550/arXiv.2306.14898

  9. [9]

    Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W

    Andy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W. Lin, Eliot Jones, Gashon Hussein, Samantha Liu, Donovan Jasper, Pura Peetathawatchai, Ari Glenn, Vikram Sivashankar, Daniel Zamoshchin, Leo Glikbarg, Derek Askaryar, Mike Yang, Teddy Zhang, Rishi Alluri, Nathan Tran, Rinnara Sangpisit, Polycarpos Yiorkadjis, Kenny Osele, Gautham ...