{"id":"8afb6fd9-43f6-4f88-8687-e236ea71ba9a","arxiv_id":"2605.11504","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Static CTF benchmarks for LLM agents are contamination-prone; CTFusion evaluates agents on live CTFs via an MCP server on CTFd with per-agent isolation.","lead":"CTFusion is a live Capture-The-Flag evaluation framework that scores LLM cybersecurity agents without reusing old challenges. It aims to stop data contamination and web-search cheating that can inflate scores on static CTF benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Abstract-only review cannot verify the central empirical claim that CTFusion is robust while static CTF benchmarks are not; the load-bearing premise about Live-CTF isolation remains uncheckable.","rationale":"The Reader's UNVERDICTED / LOW-confidence stance is the only defensible position given an abstract-only review of an empirical systems paper. The problem framing (static CTF contamination) is plausible and the engineering response (MCP server on CTFd with first-flag forwarding) is reasonable, yet none of the quantitative evidence, baselines, or artifacts needed to underwrite the robustness claim are present. My concern is identical to the Reader's weakest_assumption: Live-CTF isolation and the operational scoring choices are load-bearing and currently uncheckable. No stronger technical objection (e.g., an internal inconsistency or circular construction) can be raised from the abstract. Therefore the verdict stays UNVERDICTED; agreement with the Reader is full. The concrete test supplies a minimal, falsifiable check that would become available once the full paper or code is released.","tokens_in":2101,"tokens_out":571,"duration_ms":4426,"concrete_test":"Obtain the full paper (or the promised open-source release) and re-run the three-LLM / two-agent suite on at least one of the five Live CTFs after deliberately injecting a public write-up of one challenge into the agents' tool context; if success rates on the contaminated challenge rise substantially while CTFusion's isolation metrics remain unchanged, the robustness claim is supported; if both static and live scores inflate equally, the isolation premise fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim is that existing CTF benchmarks are unreliable due to contamination/cheating (illustrated by web-search integration) while CTFusion provides a robust alternative via Live CTFs, per-agent independence under one team account, and first-correct-flag forwarding. That claim rests on two uninspectable premises: (1) the five Live CTFs stayed free of public leakage for the evaluation window, and (2) the operational choices (single-team account + first-flag forwarding) isolate agents without introducing new scoring bias or incomplete coverage of multi-agent behavior. Because only the abstract is available, neither the experimental protocols, quantitative results, baselines, nor any code/data package can be examined; the open-source claim is therefore also unverifiable. The Reader correctly flags this as the weakest assumption. No internal contradiction appears in the abstract, but the central empirical contrast cannot be confirmed or falsified from the given text alone. This is an information-sufficiency concern, not a soundness flaw.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper claims that existing Capture-The-Flag (CTF) benchmarks for LLM-based cybersecurity agents are unreliable because they reuse published challenges, enabling data contamination and cheating (illustrated by integrating web search into an agent). It introduces CTFusion, a streaming Live-CTF evaluation framework implemented as an MCP server on CTFd, designed to preserve per-agent independence under a single team account and to reduce competition effects by forwarding only the first correct flag per challenge. Experiments with three LLMs, two agents, and five Live CTFs are said to show that static CTF benchmarks can be unreliable while CTFusion is a robust alternative; the system is released as open source.","tokens_in":2289,"tokens_out":847,"duration_ms":26547,"significance":"If substantiated, CTFusion would address a timely and practical problem: contamination and tool-enabled cheating undermine static CTF benchmarks just as LLM agents are being applied to security tasks. Strengths claimed in the abstract include an open-source MCP/CTFd implementation with broad event and agent applicability, an explicit practical demonstration of contamination via web search, and an operational design aimed at Live-CTF isolation. A validated, platform-compatible streaming evaluation method would be useful to the agent-evaluation and cybersecurity communities. Significance remains conditional on experimental evidence that is not inspectable from the abstract alone.","major_comments":[{"comment":"The central empirical claim—that existing CTF benchmarks are unreliable while CTFusion is robust—rests on experiments with three LLMs, two agents, and five Live CTFs. From the abstract alone, protocols, quantitative results, baselines, contamination measurements, statistical significance, and failure cases are unreported. Without these load-bearing details, the claimed contrast cannot be verified or falsified, so the paper’s main contribution cannot yet be assessed.","section":"Abstract"},{"comment":"The robustness claim depends on the premise that Live CTF challenges remain free of public leakage during evaluation and that single-team-account operation with first-correct-flag forwarding isolates agents without introducing scoring bias or incomplete multi-agent coverage. The abstract only sketches these operational choices; they are load-bearing axioms for preferring CTFusion over static benchmarks and require explicit justification, threat analysis, and empirical checks in the full manuscript.","section":"Abstract"},{"comment":"The practical confirmation of contamination via web-search integration is cited as evidence that existing benchmarks are unreliable. For this to support the central claim, the manuscript must specify the agent and benchmark used, success rates with versus without search, and why gains reflect contamination rather than legitimate tool use. Those details are not available here and are necessary to ground the unreliability argument.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase “preserves per-agent independence under a single team account” assumes familiarity with CTFd team mechanics; a brief clarification would help general readers.","section":"Abstract"},{"comment":"The open-source release is asserted but not linkable from the abstract; the full paper should provide a repository URL and a minimal reproducibility package (configs, agent wrappers, evaluation logs).","section":"Abstract"},{"comment":"The five Live CTFs should be characterized (categories, difficulty, duration, participant scale) so readers can judge generalizability of the robustness claim.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This is an abstract-only review; the full manuscript, code, and data were not available. The topic and framing are reasonable and timely, and there is no internal contradiction visible in the abstract, but the central empirical contrast and the Live-CTF isolation premise are currently uncheckable. I recommend obtaining the full text (and preferably the open-source artifact) before a definitive accept/revise/reject decision. Recommendation is therefore uncertain rather than reject or major_revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: this is a systems/benchmark paper that packages Live-CTF evaluation for LLM agents as an MCP server on CTFd, with a couple of operational tricks (per-agent independence under one team account, first-correct-flag forwarding) aimed at reducing contamination and competition side-effects. That combination is the actual contribution, not a new theory of agents.\n\nWhat they do well is name a real problem. Static CTF leaderboards that reuse published challenges are easy to game once web search or training data leaks in; their web-search integration anecdote is a concrete illustration, not just hand-waving. Building on CTFd and shipping as MCP is practical engineering that other groups can actually plug into. If the open-source claim holds and the five Live CTFs are documented cleanly, this is the kind of artifact people will run rather than re-implement.\n\nSoft spots, in proportion: we only have the abstract. The load-bearing claim—that existing benchmarks are unreliable while CTFusion is robust—rests on three LLMs, two agents, and five Live CTFs whose protocols, numbers, baselines, and leakage checks we cannot see. The premise that those live events stayed clean for the evaluation window, and that single-team + first-flag forwarding does not itself bias scores or hide multi-agent behavior, is stated but not independently verifiable here. That is an information gap, not an internal contradiction. Circularity risk is low; this is empirical harness work, not fitted predictions dressed up as theory. Novelty is incremental (live contests and CTFd already exist); significance is real for the agent-eval and cyber-eval subfields but not field-redefining.\n\nWho it is for: people building or ranking LLM cybersecurity agents who currently trust static CTF suites. A serious referee should see the full paper, code, and event logs. I would send it to peer review rather than desk-reject; the problem is real and the engineering response is coherent enough to deserve scrutiny. I would not cite it yet from abstract alone, and I would only bring it to reading group if someone is actively building agent eval harnesses this quarter.","headline":"Useful live-CTF harness for agent eval; central contamination claim is plausible but uncheckable from the abstract alone.","tokens_in":2940,"tokens_out":528,"would_cite":false,"duration_ms":6519,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Live CTF streaming evaluation avoids contamination and cheating that make static CTF benchmarks unreliable for LLM cybersecurity agents.","keywords":["LLM agents","cybersecurity evaluation","Capture The Flag","data contamination","Live CTF","Model Context Protocol","CTFd","agent benchmarking"],"falsifier":"Run the same three LLMs and two agents on a fresh live CTF after its challenges have already appeared in public write-ups or training corpora; if CTFusion scores then collapse to the same inflated pattern as static benchmarks, the contamination-resistance claim fails.","tokens_in":2920,"feed_emoji":"🔒","tokens_out":749,"duration_ms":8049,"temperature":0.7,"pith_summary":"The paper argues that current Capture The Flag (CTF) benchmarks for LLM-based cybersecurity agents are unreliable because they reuse published challenges, so models can draw on training data or web search instead of genuine multi-step reasoning. The authors confirm the problem by adding a web-search tool to an existing agent and observing that it can solve challenges via leakage rather than skill. They introduce CTFusion, a streaming evaluation framework that runs against live, ongoing CTF events. CTFusion keeps agents independent under a single shared team account and forwards only the first correct flag for each challenge so that competition dynamics do not distort scores. Implemented as a Model Context Protocol server on the popular CTFd platform, the system is designed to work with many CTF events and agent types. Experiments with three LLMs, two agents, and five live CTFs show that static benchmarks can mis-rank agents while CTFusion yields a more trustworthy measure of real capability. The framework is released as open source.","feed_headline":"Live CTFs fix the cheating problem in LLM cyber-agent tests","feed_subtitle":"CTFusion streams ongoing contests so agents cannot solve challenges from memory or web search","key_machinery":"CTFusion: a Model Context Protocol (MCP) server on CTFd that streams live CTF challenges, isolates agents under a single team account, and scores only the first correct flag per challenge so that competition effects and prior leakage are minimized.","core_discovery":"Existing CTF benchmarks that reuse published challenges are contaminated and cheat-prone for LLM cybersecurity agents, whereas CTFusion—a streaming live-CTF framework that preserves per-agent independence under one team account and forwards only the first correct flag—provides a robust evaluation alternative, demonstrated across three LLMs, two agents, and five live CTFs.","pith_inferences":["If live CTFs themselves begin to leak mid-event, CTFusion may need an automatic challenge-freshness filter or short-lived private instances.","The same first-correct-forwarding pattern could be reused for other multi-agent live contests (bug bounties, red-team exercises) where shared accounts are operationally convenient.","Longitudinal CTFusion scores across successive live events could become a public, contamination-resistant capability index for cybersecurity agents."],"forward_implications":["Static CTF leaderboards can no longer be trusted as evidence of agent capability once web search or memorization is available.","Agent evaluation can be repeated continuously against successive live events instead of a fixed challenge set.","Any agent that speaks the Model Context Protocol can plug into CTFd-based live CTFs without custom platform work.","Open release of the MCP server lets other labs adopt the same isolation and first-flag rules for comparable scores."],"fun_headline_variants":["CTFusion streams Live CTFs to stop LLM agents cheating on reused challenges","Static CTF benchmarks leak; CTFusion evaluates agents on ongoing contests","Live CTF streaming via CTFusion fixes contamination in cyber-agent tests","CTFusion: first-flag Live CTF method that keeps agents independent","Reused CTFs cheat-prone for LLM agents; CTFusion uses fresh Live events"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Live CTF challenges stay free of public leakage for the whole evaluation window, and the single-team-plus-first-flag design truly isolates agents without introducing new scoring bias.","fun_headline_variants_meta":{"raw":{"variants":["CTFusion streams Live CTFs to stop LLM agents cheating on reused challenges","Static CTF benchmarks leak; CTFusion evaluates agents on ongoing contests","Live CTF streaming via CTFusion fixes contamination in cyber-agent tests","CTFusion: first-flag Live CTF method that keeps agents independent","Reused CTFs cheat-prone for LLM agents; CTFusion uses fresh Live events"]},"model":"grok-4.5","effort":"low","cost_usd":0.004104,"raw_usage":{"total_tokens":1246,"prompt_tokens":747,"num_sources_used":0,"completion_tokens":100,"cost_in_usd_ticks":41040000,"prompt_tokens_details":{"text_tokens":747,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":399,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":747,"tokens_out":100,"duration_ms":4483,"temperature":1.0,"reasoning_tokens":399,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T19:03:57.516878+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same three LLMs and two agents on a fresh live CTF after its challenges have already appeared in public write-ups or training corpora; if CTFusion scores then collapse to the same inflated pattern as static benchmarks, the contamination-resistance claim fails.","supporting_citations":[],"review_version":2}