Pith. sign in

REVIEW 4 major objections 4 minor 5 references

Evaluating AI cyber capabilities with crowdsourced elicitation

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Open, prize-funded AI tracks at two Capture The Flag competitions show that crowdsourced elicitation surfaces strong offensive cyber capability: best agents ranked in the top 5% and top 10% and reliably solved challenges taking a median…

desk verdict Genuinely useful crowdsourced-elicitation data from two CTF events, but the headline one-hour capability horizon is a single-agent point estimate that needs error bars and sensitivity analysis before it can be taken as a stable result. read the letter →

arxiv 2505.19915 v2 pith:7VRO7PEP submitted 2025-05-26 cs.CR cs.AI

classification cs.CRcs.AI
keywords AIsafetycybercapabilitiesCaptureTheFlagcrowdsourcedelicitationcapabilityevaluationbountymechanismtask-completiontimehorizonLLMagents
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that open, prize-funded AI tracks at public Capture The Flag (CTF) hacking competitions are a viable way to elicit and measure offensive cyber capabilities, complementing closed in-house evaluations that have historically underestimated AI. It reports two such events: in the first, AI agents solved 19 of 20 challenges and placed in the top 5% of 400 teams; in the second, the best AI agent solved 20 of 62 diverse challenges and placed in the top 10% of 8,000 teams. Applying a recently proposed measure of AI task-completion horizons, the paper estimates that current AI agents can reliably solve cyber challenges that take a median human CTF participant about one hour. The authors read this as evidence that crowdsourced elicitation bounties can give policymakers, labs, and CTF organizers timely, inexpensive situational awareness of emerging AI capabilities.

What carries the argument

The load-bearing object is the 50%-task-completion time horizon: the amount of clock time a median human expert needs for tasks that an AI can complete half the time. In this paper it is operationalized with CTF leaderboard data: each challenge has a recorded median human solve time, and the AI's solved-and-unsolved pattern places its horizon at roughly one hour. The supporting mechanism is the competitive AI track, where prize money draws many independent teams to build and tune agents; leaderboard rank against human teams converts raw solves into a policy-interpretable capability statement.

What would settle it

For all 62 Cyber Apocalypse challenges, tabulate the median human solve time and the best agent's success, then estimate the 50% completion horizon within each challenge category such as crypto, web, pwn, and forensics; if the horizon for interactive categories is far below one hour while offline categories sit near one hour, the paper's single one-hour number is a category artifact rather than a general capability boundary.

Watch

Extended reading notes

Core claim

The paper's central claim is that crowdsourcing the elicitation process—letting many independent teams race to extract maximum task-specific performance from AI agents—produces capability signals that in-house evaluations miss. In a three-day competition against 400 teams, the strongest AI agents solved 19 of 20 offline cryptographic and reverse-engineering challenges, matching the top human teams and earning bounties; in a larger, more diverse 62-challenge event, the best agent solved 20 challenges and outperformed 90% of human teams. Using the same 50%-task-completion time horizon as prior long-task evaluations, the paper finds that AI agents can reliably solve challenges requiring one hour or less of a median human CTF participant's time, with the estimate stable across definitions of which humans count as experts. The paper concludes that open-market elicitation bounties are a practical mechanism for tracking emerging offensive cyber capabilities.

Load-bearing premise

The one-hour estimate rests on the assumption that the recorded median human solve time for each challenge is a faithful measure of expert effort, and that the 20 challenges the AI solved are representative of all challenges at that difficulty level.

Editorial extensions

If this is right

  • Small prize pools—$7,500 in the pilot—can draw enough independent teams to nearly saturate a 20-challenge CTF, so open elicitation is a low-cost addition to existing events.
  • A public AI-versus-humans leaderboard converts raw model capability into a concrete percentile statement, which is more legible to policymakers than benchmark scores.
  • The one-hour horizon gives non-experts a shared scale: an AI agent today can be expected to finish a hacking task that would occupy a median expert for about an hour.
  • Running such tracks at the hundreds of CTFs hosted yearly could produce continuous, decentralized monitoring of AI cyber capability.
  • The large-event data show AI performance drops sharply when challenges require interacting with external machines, so future tracks should deliberately include dynamic, networked tasks before generalizing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's implicit stronger claim is that competition in an open market extracts more capability than any single team can; a direct test would be to have an in-house team apply its own best elicitation to the same 62 challenges and compare solved counts with the crowdsourced leaderboard.
  • The near-saturation of the offline crypto and reverse-engineering event combined with only a 20-of-62 score in the diverse event suggests the one-hour horizon is really a statement about access rather than reasoning; a testable version would compare horizons for offline and online challenge categories separately.
  • If bounties reliably surface capabilities before safety teams expect them, then public AI CTF leaderboards function as early warnings but also as capability releases; an evaluation policy would need to decide whether to embargo findings before publication.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper reports on two crowdsourced AI CTF tracks organized by Palisade Research with Hack The Box. In the first event (AI vs. Humans, 403 teams), seven AI agents solved most of 20 challenges, with the best ranking 20th overall. In the second event (Cyber Apocalypse, roughly 8,000 human teams), four AI agents from two teams participated; the best agent, CAI, solved 20 of 62 challenges and ranked 859th, which the paper reports as top-10%. The authors use the Cyber Apocalypse data to estimate METR's 50%-task-completion time horizon and conclude that AI agents can reliably solve cyber challenges requiring about one hour of effort from a median human CTF participant. They propose crowdsourced elicitation bounties as a complement to in-house elicitation.

Significance. The paper's main empirical contribution is a real-world, competitive measurement of AI cyber capabilities against a large human baseline. The leaderboard data are direct measurements, and the authors make the data public, which is a genuine strength. The observation that a purpose-built agent can outperform the large majority of human CTF teams, and that several independently designed agents solved 19/20 challenges in a calibrated pilot event, is valuable evidence that open-market elicitation can surface capable systems at relatively low cost. However, the headline capability horizon ('reliably solve about one hour') is not yet supported at the level of precision implied by the abstract and conclusion. The quantitative analysis lacks uncertainty quantification, and the sensitivity analysis covers only one dimension of the estimate.

major comments (4)
  1. [Section 5, Appendix C, Figure 4] The central quantitative claim — that AI agents can reliably solve challenges requiring one hour or less of human effort — is a point estimate derived from a single agent (CAI) solving 20 of 62 Cyber Apocalypse challenges. No confidence interval, bootstrap, or other uncertainty quantification is reported for the 50%-task-completion time horizon, and Figure 4 shows no error bars. Because the abstract and conclusion present this figure as a stable result, the paper should report the distribution of human solve times for solved versus unsolved challenges and provide a bootstrap over challenges (and ideally over agents) to bound the estimate.
  2. [Section 5, Figure 5] The sensitivity analysis varies only which percentile of human teams is treated as expert; it does not vary the set of 62 challenges, the CAI success pattern, or the underlying human solve-time measurements. The paper never reports the human solve-time distribution for the 42 unsolved challenges, so a reader cannot tell whether the one-hour threshold is anchored by a dense set of challenges near one hour or by a sparse region of the difficulty distribution. Please add a sensitivity check that excludes challenges near the threshold, recomputes the horizon on random challenge subsets, and reports the human solve-time distribution for both solved and unsolved challenges.
  3. [Table 2, Section 5] The capability claim rests on a single purpose-built agent: CAI, which required roughly 500 dev-hours of engineering (Appendix B.1), while the other three agents solved 2–5 challenges. The phrase 'AI agents can reliably solve' overstates the evidence, which supports a claim about the best elicited agent rather than about AI agents in general. Please either restrict the wording accordingly or show agent-wise horizons and discuss the heterogeneity across agents.
  4. [Section 3, Section 6] The AI vs. Humans event was calibrated so that the authors' own React&Plan agent would solve about 50% of the tasks, and four agents subsequently solved 19/20. The paper should state explicitly that this event does not provide an uncalibrated measure of general AI capability. Its value is as a demonstration of the elicitation mechanism and of human-competitive speed, not as independent support for the one-hour horizon.
minor comments (4)
  1. [Abstract, Table 2] The abstract reports 'top-10%' for Cyber Apocalypse, but CAI's rank of 859 among 8,129 registered teams corresponds to the 10.6th percentile, and 859th among 3,994 teams that solved at least one challenge corresponds to the 21.5th percentile. Please clarify which denominator is used and whether 'top-10%' is rounded.
  2. [References] The reference to 'Project Zero (2024)' is incomplete; it should include the full title, URL, and year of the Project Naptime post so that readers can locate the source.
  3. [Figures 4 and 5] The figure captions should state exactly what quantity is plotted (for example, median solve time, 50% completion horizon, and which human-team percentile is used) and whether any bands or intervals are shown.
  4. [Appendix B.2] The quote 'I spent 17 dev-hours on agent design' is attributed only to 'the participant'; please identify the agent or anonymize the attribution consistently.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the one-hour capability horizon is a measured statistic from independent human solve-time and AI success data, not a fitted parameter renamed as a prediction.

full rationale

The paper's central quantitative claim—'AI agents can reliably solve cyber challenges requiring one hour or less of effort from a median human CTF participant'—is an empirical estimate, not a derivation from circular premises. The AI success data (CAI solving 20/62 Cyber Apocalypse challenges) and the human effort data (Hack The Box's measured median participant solve times) are collected independently and then combined according to METR's externally defined 50%-task-completion time horizon metric, as described in Appendix C: 'To estimate the human expert effort equivalent to current AI capabilities we follow (Kwa et al. 2025) by measuring the 50%-task-completion time horizon.' No equation in the paper defines the AI solve set in terms of the human solve times, nor are the human times fitted to make the one-hour conclusion; the conclusion is a summary statistic of the two measurements. The self-citations to Turtayev et al. (2024) and Mayoral-Vilches et al. (2025) describe agent designs and are not load-bearing for the one-hour horizon. The calibration of AI vs. Humans challenges using the authors' own React&Plan agent affected that event's difficulty, but the one-hour analysis explicitly uses Cyber Apocalypse, which was not calibrated to AI ('As this competition's challenges were not saturated by AI, the resulting data allowed us to calculate the 50%-completion-time horizon'). Concerns about the lack of confidence intervals or the representativeness of the 20 solved challenges are robustness or correctness issues, not circularity. The paper is self-contained as an empirical measurement against an external benchmark methodology.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claims rest on a small number of domain assumptions rather than fitted models. The main free choices are the expert percentile for the time-horizon estimate and the difficulty calibration for the first event. No invented entities.

free parameters (2)
  • Expert percentile for human solve-time calculation = top 1% of human teams (Figure 4)
    The headline one-hour horizon uses the top 1% of human teams as the expert reference. The paper explores other thresholds in Figure 5, showing the estimate varies by a factor of several; this is a modeling choice rather than a quantity fitted to data.
  • Challenge difficulty calibration target = ~50% solve rate for organizers' React&Plan agent
    In the AI versus Humans event, challenges were selected so the authors' own agent would solve about half of them, which shapes the absolute standings and makes the saturation result less surprising.
assumptions (4)
  • domain assumption CTF performance is a valid proxy for real-world offensive cyber capability
    The entire policy relevance of the paper depends on this mapping, but it is asserted implicitly in the Introduction and never tested.
  • domain assumption Hack The Box's median solve time is a faithful measure of human expert effort
    Used in Appendix C to compute the 50% completion horizon; the paper adopts it without validating it against controlled timing or expert baselines.
  • domain assumption The CAI agent's solve record is representative of what autonomous AI agents can do
    Section 5 and Figure 4 base the one-hour claim on a single agent at one event, yet the conclusion generalizes to AI agents.
  • domain assumption METR's methodology transfers from software engineering tasks to CTF challenges
    The paper applies Kwa et al. (2025) to a new domain; the transfer is plausible but the details of the transfer are not discussed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating AI cyber capabilities with crowdsourced elicitation." pith.science (2026). https://pith.science/paper/7VRO7PEP

@misc{pith2026250519915,
  author       = {Pith},
  title        = {Pith review of: Evaluating AI cyber capabilities with crowdsourced elicitation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7VRO7PEP}},
  note         = {Machine review of arXiv:2505.19915}
}
abstract

As AI systems become increasingly capable, understanding their offensive cyber potential is critical for informed governance and responsible deployment. However, it's hard to accurately bound their capabilities, and some prior evaluations dramatically underestimated them. The art of extracting maximum task-specific performance from AIs is called "AI elicitation", and today's safety organizations typically conduct it in-house. In this paper, we explore crowdsourcing elicitation efforts as an alternative to in-house elicitation work. We host open-access AI tracks at two Capture The Flag (CTF) competitions: AI vs. Humans (400 teams) and Cyber Apocalypse (8000 teams). The AI teams achieve outstanding performance at both events, ranking top-5% and top-10% respectively for a total of \$7500 in bounties. This impressive performance suggests that open-market elicitation may offer an effective complement to in-house elicitation. We propose elicitation bounties as a practical mechanism for maintaining timely, cost-effective situational awareness of emerging AI capabilities. Another advantage of open elicitations is the option to collect human performance data at scale. Applying METR's methodology, we found that AI agents can reliably solve cyber challenges requiring one hour or less of effort from a median human CTF participant.

Figures

Figures reproduced from arXiv: 2505.19915 by the authors.

Figure 1
Figure 1. Challenges solved over time for all teams, [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Challenges solved over time for top teams, [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Challenges solved over time for top teams, [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: 50%-task-completion time horizon, Cyber Apocalypse (top 1% human teams) 6 Conclusion Accurately assessing the offensive capabilities of AI systems remains a major challenge for policymakers and researchers. Relying solely on evaluations conducted by a single team may b…
Figure 5
Figure 5. Figure 5: 50%-task-completion time horizon estimates, depending on which percentage of top [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

5 extracted references · 1 canonical work pages

  1. [3]

    Project Zero

    https://arxiv.org/abs/ 2504.06017. Project Zero

  2. [4]

    Yang, John, Akshara Prabhakar, Shunyu Yao, Kexin Pei, and Karthik R Narasimhan

    https://arxiv.org/abs/2412.02776. Yang, John, Akshara Prabhakar, Shunyu Yao, Kexin Pei, and Karthik R Narasimhan

  3. [5]

    B.1 CAI Used a custom harness design they spent about 500 dev-hours on, see (Mayoral-Vilches et al

    B Submitted agent designs Below are some of the agent designs used by AI teams. B.1 CAI Used a custom harness design they spent about 500 dev-hours on, see (Mayoral-Vilches et al. 2025). B.2 Imperturbable From the participant: I spent 17 dev-hours on agent design I used EnIGMA (with some modifications) and Claude Code, with different prompts for rev/crypt...

  4. [2024]

    Kwa, Thomas, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, et al

    https://arxiv.org/abs/2404.13161. Kwa, Thomas, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, et al

  5. [2025]

    org/abs/2503.14499

    https://arxiv. org/abs/2503.14499. Mayoral-Vilches, Víctor, Luis Javier Navarrete-Lozano, María Sanz-Gómez, Lidia Salas Espejo, Martiño Crespo-Álvarez, Francisco Oca-Gonzalez, Francesco Balassone, et al

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.