{"id":"92097eb9-e134-4915-b6a3-959942bb089c","arxiv_id":"2607.11698","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A falsifiable agent-vs-agent loop discovers a frozen Vulnerability Concept Graph that beats frozen red-team baselines by 14.2 points and transfers across scenarios, channels, and models.","lead":"AHA turns red-teaming of production coding agents into a falsifiable research loop that stores reusable vulnerability concepts, not just attack payloads. Safety teams get an auditable graph they can re-run after patches without re-searching.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The 14.2-point frozen-VCG gain may partly reflect research-model asymmetry and judge/falsifier residual dependence rather than pure concept reusability.","rationale":"The reader correctly identifies the weakest assumption as judge+falsifier reliability separating mechanism-level breaks from artifacts, and correctly keeps CONDITIONAL rather than ACCEPT given model asymmetry, sparse AgentDyn, and residual harness dependence. My stress test does not invent a new flaw; it sharpens the same soft spot into the single most load-bearing threat to the strongest claim: the 14.2 pp frozen-artifact advantage is the number that must be method-pure for “reusable vulnerability knowledge” to be the right unit. The paper’s protocol (frozen single-shot, ablations, transfer, public code) is strong and the shared claimed-authorization core plus cross-scenario transfer are real positive evidence. But the incomplete AHA@Qwen control and the still-judge-mediated success metric mean the headline gain could partly be research-model strength or residual scoring artifacts rather than VCG reusability alone. That does not overturn the claim as stated under the paper’s protocol, so the verdict stays CONDITIONAL; it does not warrant REJECT. Agreement with the reader is full on the weakest assumption and on the overall posture.","tokens_in":49293,"tokens_out":793,"duration_ms":7193,"concrete_test":"Re-run the full frozen single-shot protocol of Table 2 with AHA discovery forced onto the same Qwen-3.7-Max attacker used by the baselines (extend Table 3 to all three scenarios × both agents × three victims), and recompute the overall Avg ASR gap. If the gap falls below ~5–7 pp (or loses leadership on Claude Code AgentHazard/AgentDyn), the 14.2-point method claim is not model-pure. Optionally, re-score a stratified sample of AHA held-out “breaks” with an independent human or second-judge panel against the committed falsifier; if >15–20% are reclassified non-mechanism, residual judge dependence is material.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that a frozen VCG from AHA’s falsifiable loop is reusable without further search and beats the strongest frozen discovery baseline by 14.2 pp under a shared single-shot protocol (Table 2; RQ2), while encoding shared mechanisms (RQ1) that transfer (RQ4). The load-bearing soft spot is that this gain is not cleanly isolated from two confounds the paper itself surfaces. First, AHA’s discovery loop runs on the host model (Claude-4.8-Opus / GPT-5.5), while T-MAP*, IterInject*, and AutoRISE* payload generation use compliant Qwen-3.7-Max because host models refuse the same prompts (§4.4, B.6). The AHA@Qwen isolation (Table 3) is only on Deepseek×Claude Code and only vs T-MAP*/IterInject*, not the full 18-setting average that produces the 14.2-point headline. Second, held-out ASR and transfer ASR still rest on scenario judges (AgentHazard LLM score≥7, AgentDyn AgentDojo security, DTap backend predicates) plus the committed falsifier (§3.4–3.5, C.1). RQ3 shows removing the falsifier/memory/critic leaves discovery ASR high while cutting held-out survival, which supports the safeguards—but does not prove residual judge-gamed or harness-specific breaks are fully screened once the VCG is frozen. If either confound drives a large share of the 14.2 pp, the claim that the VCG itself is the reusable unit of knowledge is weaker than stated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes AHA, an autoresearch loop in which one agentic research environment discovers reusable vulnerability knowledge about production-style tool-using agents (Claude Code, Codex). The loop commits a vulnerability hypothesis and falsifier before attack design, executes scenario-valid payloads in a sandboxed harness, reflects on trajectories, and promotes confirmed findings into a Vulnerability Concept Graph (VCG) whose entries package claim, enabling condition, falsifier, template, transfer prediction, and evidence. On AgentHazard, AgentDyn, and DTap, frozen VCGs are evaluated single-shot on held-out instances against frozen discovery baselines (T-MAP*, IterInject*, AutoRISE*) and the benchmark original attack. The authors report a 14.2-point overall ASR gain over the strongest frozen baseline, a recurring claimed-authorization core across models and agents, ablations of falsifier/memory/critic, and cross-scenario and cross-victim transfer.","tokens_in":49718,"tokens_out":1065,"duration_ms":15786,"significance":"If the results hold under cleaner isolation, this is a meaningful reframing of agent red-teaming: from perishable payloads/archives/programs toward auditable, mechanism-level concepts that can be inspected, patched against, and re-run after model or product updates. The experimental protocol is unusually careful for the area—fixed discovery/held-out splits, frozen artifacts, one attack per held-out instance, shared evaluator/sandbox/judge, and transfer tests that go beyond leaderboard ASR. The VCG design, negative-result bookkeeping, and bring-your-own build/import scenario workflows are practically useful artifacts for production safety teams. Code release further strengthens the contribution. The work is dual-use but oriented toward defensive triage rather than exploit dissemination.","major_comments":[{"comment":"The headline 14.2 pp frozen-artifact gain (abstract; §5.2; Table 2) is not cleanly isolated from research-model asymmetry. AHA discovery runs on host models (Claude-4.8-Opus / GPT-5.5), while T-MAP*, IterInject*, and AutoRISE* payload generation use compliant Qwen-3.7-Max because host models refuse the same prompts (§4.4, B.6). The AHA@Qwen control (Table 3) is limited to Deepseek×Claude Code and only vs T-MAP*/IterInject*, not the full 18-setting average that produces the 14.2-point claim. AutoRISE is further hybrid (coding agent on the host model). Either expand the matched-model isolation across scenarios/victims, or revise the abstract/intro claim to a more carefully scoped comparison that does not present 14.2 pp as pure concept-reusability advantage.","section":null},{"comment":"Held-out and transfer ASR still rest on scenario judges plus the committed falsifier (§3.4–3.5, C.1). AgentHazard’s break rule is an LLM trajectory judge with harmful flag and score ≥7; AgentDyn/DTap use verifiable checks. RQ3 (Table 4) shows that removing falsifier/memory/critic leaves discovery ASR high while cutting held-out survival—good evidence for the safeguards—but does not quantify residual judge-gamed, off-target, or harness-specific breaks that survive promotion into the frozen VCG. For the claim that ASR measures reusable vulnerability knowledge rather than optimizer/judge artifacts, add either (i) a human audit of a stratified sample of held-out ‘breaks’ against the committed enabling condition, or (ii) a sensitivity analysis under stricter AgentHazard thresholds / alternative judges, and discuss residual risk explicitly in §5–6.","section":null},{"comment":"Promotion and search knobs are free parameters without sensitivity (Eq. 2; Table 8: n_conf ≥3, confirmation rate ≥0.6, confidence ≥0.6, outer cap 100, mode availability rule). Because the counted VCG is the deployable artifact, the main results could shift under nearby thresholds. Report at least a small sensitivity sweep on n_conf and confirmation rate for one representative setting (e.g., Claude Code×Minimax×AgentHazard), or justify the thresholds as fixed a priori and show that effective concept count / held-out ASR is stable in a neighborhood of the chosen values.","section":null},{"comment":"Codex×DTap evaluation is partially compromised by tool-routing instability (footnote in §5.2: non-OpenAI backends and type:namespace MCP encoding; openai/codex#26234). Multi-tool indirect attacks fail in unstable mode with unsupported-tool errors, which depresses Codex DTap numbers and muddies cross-agent comparison on that scenario. Either restrict the Codex DTap claim to settings where tool calls are reliable, re-run with a fixed routing fix, or move Codex DTap to appendix with an explicit reliability filter so Table 2 averages are not diluted by harness noise.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful thing here is not another jailbreak optimizer. It is a change in the retained artifact: a Vulnerability Concept Graph with a pre-committed falsifier, evidence-gated promotion, and frozen single-shot reuse. That is the right unit for production-agent safety work, and the paper actually measures it.\n\nWhat is new is the full loop plus the evaluation discipline. They run black-box discovery on Claude Code and Codex across AgentHazard, AgentDyn, and DTap, freeze the VCG, and instantiate exactly one attack per held-out instance with no test-time search. Table 2 and the RQ1–RQ4 package are the real contribution: a shared claimed-authorization core across models and agents, 14.2-point average gain over the strongest frozen discovery baseline, ablations showing falsifier/memory/critic matter for held-out survival rather than discovery ASR, and transfer across scenarios and direct/indirect channels. Code is public. The experimental design is careful for this area—fixed splits, shared sandbox/judge budget, and explicit separation of discovery from evaluation.\n\nThe soft spots are real but proportionate. Baseline payload generation uses Qwen because host models refuse the same prompts; AHA@Qwen (Table 3) is only a partial isolation on Deepseek×Claude Code, not the full 18-setting average behind the headline number. AgentDyn absolute ASRs stay low, and held-out/transfer numbers still rest on scenario judges plus the falsifier. RQ3 supports the safeguards without fully proving residual judge or harness artifacts are gone. Those are standard red-team evaluation limits, not circular math or invented targets. Free parameters (promotion thresholds, mode rules) are disclosed and not hidden.\n\nThis is for people building or securing tool-using agents who need auditable, re-checkable failure mechanisms rather than leaderboard payloads. I would bring it to reading group, cite the protocol and VCG idea, and send it to peer review. Tighten the same-attacker-model controls and the judge-dependence discussion; the central claim still holds under the protocol they actually ran.","headline":"Solid systems paper: falsifiable autoresearch that freezes a reusable VCG beats frozen baselines by ~14pp under a careful single-shot protocol; the main soft spot is incomplete isolation of research-model strength, not a broken claim.","tokens_in":50328,"tokens_out":525,"would_cite":true,"duration_ms":6708,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A falsifiable autoresearch loop turns production-agent red-team runs into a frozen Vulnerability Concept Graph that stays reusable without further search and beats frozen attack archives by 14.2 points.","keywords":["LLM agents","automated red-teaming","vulnerability concept graph","falsifiable discovery","indirect prompt injection","production agent safety","autoresearch","attack transfer"],"falsifier":"Freeze a VCG discovered on one scenario or victim model, deploy it single-shot with no further search on held-out instances from a different scenario or model (including direct-to-indirect channel shifts), and check whether held-out attack success collapses to baseline levels or the claimed-authorization core fails to recur across independent victim agents.","tokens_in":50175,"feed_emoji":"🔓","tokens_out":1059,"duration_ms":13378,"temperature":0.7,"pith_summary":"Production coding agents act on untrusted files, tools, and workspace state, so a safety failure is a real action rather than a bad sentence. Most automatic red-teaming optimizes attack success and freezes payloads, archives, or programs, which record where a break landed but not the enabling condition that made the agent trajectory unsafe. This paper reframes the problem as autoresearch: one research agent states a vulnerability hypothesis and falsifier before designing an attack, runs it in a sandbox, reflects on the trajectory, and promotes only confirmed findings into a Vulnerability Concept Graph. Across Claude Code and Codex on direct and indirect scenarios, the frozen graph needs no further search at test time, outperforms the strongest frozen discovery baseline by 14.2 percentage points under a shared single-shot protocol, and transfers across scenarios, attack channels, and victim models. The intended payoff is cumulative, auditable safety knowledge that a team can inspect, patch against, and re-check after model or product updates.","feed_headline":"Frozen vulnerability concepts beat attack archives by 14 points","feed_subtitle":"A falsifiable agent loop turns red-team runs into reusable maps of how coding agents fail.","key_machinery":"The Vulnerability Concept Graph (VCG), filled by AHA’s falsifiable discovery loop: commit a hypothesis and falsifier before any attack, instantiate a scenario-valid payload, execute in a sandbox, adjudicate the trajectory against the falsifier, and promote only concepts that meet a multi-confirmation evidence rule into a reusable, auditable library.","core_discovery":"The authors claim that black-box red-teaming of tool-using production agents can be organized as a falsifiable discovery loop whose lasting product is not a payload archive but a Vulnerability Concept Graph: each concept links an attacker-facing surface to an unsafe trajectory through a claim, enabling condition, committed falsifier, transfer prediction, and evidence. A graph that passes their promotion rule can be frozen and reused single-shot, revealing a shared claimed-authorization core across models and agents and outperforming the strongest frozen discovery baseline by 14.2 percentage points on held-out instances.","pith_inferences":["If the shared claimed-authorization core is truly agent-level, product teams might prioritize provenance and ownership checks over content filters alone when hardening coding agents.","The same hypothesize–falsify–promote contract could be applied to browser, mobile, or multimodal agents wherever an attack surface, harness, trajectory, and judge exist.","A public, versioned VCG across vendors could become a common language for incident response, analogous to how CVE entries accumulate software-vulnerability knowledge.","Judge gaming remains the practical bottleneck: any improvement in trajectory judges would tighten the evidence rule and change which concepts survive promotion."],"forward_implications":["A frozen VCG can be re-run after a safety patch or model swap to check whether a named enabling condition still produces an unsafe trajectory, without launching a new search.","Safety teams can treat each concept’s enabling condition as a concrete lever for policy, tool-boundary, or workflow changes rather than chasing perishable payloads.","Concepts that transfer from direct multi-turn attacks to indirect tool-mediated injection (and the reverse) locate the failure in agent trajectory structure, not in stored attack strings.","Build and import workflows can fold new free-text concerns or existing benchmarks into the same loop, so new threats extend one shared concept library instead of spawning one-off silos.","Negative results (falsified framings kept as non-mechanisms) become part of the audit trail so later search does not re-explore inert patterns."],"fun_headline_variants":["Frozen VCG beats attack archives by 14 points","Vulnerability concepts transfer across agents and models","Falsifiable loop freezes reusable agent failure maps","Shared claimed-authorization core spans coding agents","Single-shot frozen VCG tops discovery baselines 14.2 pts"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That the scenario judges plus the pre-committed falsifier reliably separate true mechanism-level agent failures from judge-gamed, off-target, or harness-specific breaks, so held-out and transfer scores measure reusable vulnerability knowledge rather than artifacts of the search or the scorer.","fun_headline_variants_meta":{"raw":{"variants":["Frozen VCG beats attack archives by 14 points","Vulnerability concepts transfer across agents and models","Falsifiable loop freezes reusable agent failure maps","Shared claimed-authorization core spans coding agents","Single-shot frozen VCG tops discovery baselines 14.2 pts"]},"model":"grok-4.5","effort":"low","cost_usd":0.003236,"raw_usage":{"total_tokens":1150,"prompt_tokens":868,"num_sources_used":0,"completion_tokens":57,"cost_in_usd_ticks":32360000,"prompt_tokens_details":{"text_tokens":868,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":225,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":868,"tokens_out":57,"duration_ms":2332,"temperature":1.0,"reasoning_tokens":225,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T03:46:40.483825+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Freeze a VCG discovered on one scenario or victim model, deploy it single-shot with no further search on held-out instances from a different scenario or model (including direct-to-indirect channel shifts), and check whether held-out attack success collapses to baseline levels or the claimed-authorization core fails to recur across independent victim agents.","supporting_citations":[],"review_version":1}