{"id":"b9c6e6ac-d3c0-4c1f-a6d1-8763dc2eda70","arxiv_id":"2608.03499","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A benchmark and Docker sandbox that tests whether owner-scoped AI agents can collaborate on real tasks without being manipulated into privacy leaks, poisoned evidence, or invalid approvals.","lead":"WeClawArena is a new testing ground where AI agents working for different people must finish shared tasks while each one can only see its own files, tools, and rules, and where attackers try to trick them into leaking data or breaking approvals. It measures task success and attack harm separately, and keeps a log of every message and action so failures can be diagnosed.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing soft spot is the ASR judge: it consumes author-authored 'expected exposure path' and 'evidence hints' (Fig. 8), so moderate human agreement (κ=0.65–0.68) does not yet rule out hint-driven over-attribution. A hint-redaction rerun would settle it.","rationale":"The reader's weakest assumption matches my read: the ASR judge's measurement validity is the load-bearing pillar for the empirical separation claim. The paper has strong internal support—deterministic TSR, exact denominators, a second judge with high ranking concordance, low benign FPR, and unusually candid Appendix H limitations—so this is not a rejection-level flaw. But the hints issue is not resolved by inter-judge agreement because both judges consume the same potentially priming fields, and the human pilot is small and co-authored. κ=0.65–0.68, while 'substantial,' leaves room for per-cell ASR disagreement that could shift the headline model ordering. The proposed hint-redaction test is feasible with the released judge prompts and evidence packets and would directly settle whether reported ASR tracks final harm or designer expectations. Since the reader's CONDITIONAL verdict already reflects exactly this uncertainty, I recommend no change to the verdict.","tokens_in":27845,"tokens_out":5253,"duration_ms":61036,"concrete_test":"Rerun the two ASR judges (GPT-5.2 and Opus 4.7) on a stratified sample of at least 300 rows from ASR-MAIN-SIX-DOMAIN with the 'expected exposure path' and 'evidence hints' fields removed from the evidence packet, while keeping transcript messages, tool calls, observations, final state, and the attack-vector label. Pre-register a threshold, e.g., per-model macro-vector 1-ASR shift >5 pp or any model rank swap, as evidence that reported ASR is partially driven by hints. Additionally, have non-author experts label 200 of the same redacted packets; if human-judge agreement stays near 0.65–0.68 and the model ranking is unchanged, the concern does not land. If ASR drops materially or the ranking changes, the headline resistance numbers should be re-reported with hint-free judging.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claims—that task utility and final harm are separable and that model attack-resistance differs sharply (Opus 4.7 at 97.4% 1-ASR vs. 51.1–64.0% for open-weight models)—depend on ASR measuring real final harm. The judge's evidence packet (Figure 8, Appendix E) includes the attack designers' 'expected exposure path' and 'evidence hints' for every attacked row, and the judge is asked whether final harm occurred with a clear evidence link to the attack. Those fields tell the judge which trace events to treat as harmful; they are absent from benign controls, so the 0.3% benign FPR does not control for hint-driven over-attribution. The 200-row human pilot is the only direct check against that, and it is small, author-annotated, and yields κ=0.65–0.68 (Appendix H); both LLM judges and the annotators see the same hints. The second-judge sensitivity pass (Opus 4.7, inter-judge κ=0.70, ranking Spearman 1.00) is genuine support: the ranking is robust to LLM choice. But it does not test the common hint structure. If ASR is inflated by hints, the 117 task-success-plus-attack-success rows and the model resistance ranking shift; the separation claim itself is not falsified, but a key measurement pillar is uncertain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"WeClawArena introduces a benchmark and Docker-based sandbox for multi-party owned-agent collaboration over personal workspaces. It contains 124 base tasks expanded into 620 scenario variants across six domains, with each base task paired with one benign control and four attack-vector variants (collaboration, security, privacy, governance). The paper's central methodological claim is that task utility and final attack harm are separable and measurable from bounded runtime evidence: utility is scored by deterministic task verifiers, while attack success rate (ASR) is judged post-hoc by an LLM from bounded evidence packets. Experiments across eight models report TSR and 1−ASR separately, with Claude Opus 4.7 as the most attack-resistant model. The paper includes a second-judge sensitivity pass, a 200-row human-annotated pilot, Wilson confidence intervals, and a detailed artifact-release plan.","tokens_in":28123,"tokens_out":10761,"duration_ms":120454,"significance":"This is a timely and potentially valuable benchmark: it targets a realistic deployment setting—persistent personal agents collaborating across owner-scoped workspaces—that is not covered by existing tool-use, multi-agent, or privacy benchmarks. The design has real strengths: utility and harm are explicitly separated; the runtime records gateway-mediated evidence; attack variants are matched to benign controls; and the paper is unusually transparent, with exact denominator tables, Wilson intervals, an independent LLM-judge sensitivity pass, and an explicit audit-claim boundary in Appendix A. The 117 task-success-plus-attack-success rows and the per-surface heterogeneity results, if reproducible, support the conceptual claim that utility and harm must be measured separately. However, the headline ASR numbers—and therefore the model-resistance ranking and the utility/harm separation evidence—rest on a single LLM judge whose evidence packets contain the attack designers' own 'expected exposure path' and 'evidence hints'. The current validation does not yet rule out hint-driven over-attribution, and there is a direct numerical inconsistency between the reported ASR pool sizes and the ca","major_comments":[{"comment":"The ASR denominator arithmetic is internally inconsistent. Appendix G states that ASR-MAIN-SIX-DOMAIN contains 3,743 judged attack-vector rows, with per-vector denominators 939 (collaboration), 922 (security), 943 (privacy), and 939 (governance). Table 11, described as the canonical denominator table for that pool, sums to 3,302 rows (collaboration 930, security 778, privacy 799, governance 795), and Table 3's per-model ASR denominators also sum to 3,302. The ASR-success numerators sum to 1,152 in both places. The numbers only close if the 441 rows present in Appendix G but absent from Table 11 are all non-successes, which would change the overall row-micro ASR from the reported 30.8% to 34.9%. Please reconcile these counts and state exactly which rows belong to ASR-MAIN-SIX-DOMAIN, ASR-SENSITIVITY-SIX-DOMAIN, and Table 11. This is load-bearing: every ASR percentage, the failure-analysis","section":"Appendix G / Table 3 / Table 11"},{"comment":"The ASR judge is the load-bearing measurement for the separation claim and the model-resistance ranking, but the validation does not rule out hint-driven over-attribution. Figure 8 shows that the judge's evidence packet includes the attack recipe's 'expected exposure path' and 'evidence hints' for every attacked row, and Appendix H states that the two human annotators saw the same packets. Inter-judge κ=0.70 and human-judge κ=0.65–0.68 therefore share the same hint structure, and the 0.3–0.7% benign FPR is not an adequate control because benign rows contain no attack metadata or hints. I recommend a hint-redaction rerun: judge the same attacked rows with 'expected exposure path' and 'evidence hints' removed, and compare per-cell ASR; alternatively, have experts label a redacted subset. If ASR falls substantially, the headline 1−ASR values and the 117 task-success-plus-harm rows would shi","section":"Appendix E / Appendix H / Figure 8"},{"comment":"Even after the denominator reconciliation, the model-level comparison must address non-random missingness. Several models have large gaps between nominal attack rows (496 per model) and judged rows: Claude Sonnet 4.5 has 293/496, and Qwen3 235B has 392/496, with other models in the 434–440 range. If unjudged rows correlate with task difficulty or judge failure, macro-vector resistance is biased. Please report the reason each row is absent (e.g., missing evidence, judge error, runtime failure) and add a sensitivity analysis that imputes missing rows as all-success/all-failure or restricts the ranking to the fully covered common subset. Relatedly, Appendix A excludes runs with gateway bypasses from valid evidence; the paper should report how many runs were excluded on that basis so readers can assess selection effects.","section":"Table 3 / Appendix I.1 / Appendix A"}],"minor_comments":[{"comment":"The caption should state explicitly that TSR is computed over all five variants, including attacked variants. As written, readers may misread the clinical and trading columns as benign-capability scores rather than all-variant resilience scores.","section":"Table 1 caption"},{"comment":"The 'A2A' label in the left panel is not defined in the caption. The relationship between the Agent2Agent protocol cited in Related Work and the gateway used in this paper should be clarified.","section":"Figure 1"},{"comment":"The statement that GPT-5.2's 'consistently lower ASR' means the headline numbers are a 'lower-bound on attack success' assumes that the lower-ASR judge is closer to truth. This is not established by the validation; if hint-driven over-attribution inflates ASR, the direction of the bias is the opposite. Please qualify or justify.","section":"Appendix H"},{"comment":"'WeClawArena establishes multi-party tool-use collaboration...' overstates the contribution; consider 'introduces' or 'provides' to match the benchmark framing used elsewhere.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the journal's scope and the benchmark is potentially a strong community resource. The main issue is not the conceptual design but the reliability and reporting of ASR. The numerical discrepancy between Appendix G and Table 11 must be fixed before the paper can be accepted. The human pilot is small and author-annotated; even with a redaction rerun, an external annotator or a larger panel would substantially strengthen the audit claim. I would ask the authors to treat the ASR validation as a first-class contribution rather than an appendix-level sensitivity check."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The benchmark design is the real contribution. WeClawArena gives the community a concrete substrate for testing owned agents that collaborate across private workspaces, with one benign control and four matched attack variants per base task, and with utility scored deterministically while attack success is judged separately from bounded runtime evidence. That separation, plus published denominators, Wilson intervals, and a disclosed-limitations section, is better than most benchmark papers manage. The six domains are varied enough that the utility/resistance tradeoff plots are informative. The formal model in Section 2 is standard, but the artifact is the point.\n\nThe main soft spot is the ASR judge. The evidence packet includes the attack authors' \"expected exposure path\" and \"evidence hints\" (Figure 8, Appendix E). The judge is asked to confirm final harm with a clear evidence link, but those fields essentially tell it which trace events to treat as harmful. The benign FPR of 0.3% does not control for this, because benign rows have no hints. The 200-row human pilot is small, author-annotated, and the annotators saw the same hints; kappa 0.65-0.68 is moderate. The second-judge pass (Opus 4.7) is genuine support that the model ranking is not a fluke of LLM choice (Spearman 1.0), but it does not test the common hint structure. So the headline ASR numbers—the 97.4% vs 51-64% resistance spread, the 117 task-success-plus-attack rows—are more fragile than the table formatting suggests. A hint-redaction rerun or an annotation study without hints would settle it. This is not fatal: the utility/ASR separation claim holds as a design, and the authors disclose the issue openly.\n\nMinor points: a few headline cells have partial coverage (Kimi K2.5, Qwen3 235B), flagged with daggers; artifact links couldn't be verified in this pass. Both are minor and checkable.\n\nWho reads it: people building personal-agent platforms, multi-agent security evaluators, and anyone designing benchmarks with LLM-judged harm. It deserves a serious referee, not a desk reject. The referee should push for the hint-redaction experiment or independent no-hint annotation before the attack-resistance rankings are treated as settled; the benchmark itself will be useful either way.","headline":"A genuinely useful new benchmark for cross-user owned-agent collaboration, with a real question about whether the ASR judge measures harm or author hints — worth refereeing, and the authors are honest about most of it.","tokens_in":28726,"tokens_out":1689,"would_cite":true,"duration_ms":18001,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In human-centered agent networks, finishing a task and being attacked are independent failures, and WeClawArena measures both from bounded runtime evidence.","keywords":["human-centered agent networks","personal workspaces","multi-agent collaboration","attack success rate","auditable sandbox","governance harm","LLM-as-a-judge","privacy leakage"],"falsifier":"An audit study that removes the 'expected exposure path' and 'evidence hints' from the evidence packets and has a judge or blinded human annotators re-label the 200-row pilot: if ASR drops substantially or agreement collapses, the original judge was leaning on the attack description rather than on recorded evidence of final harm.","tokens_in":27635,"feed_emoji":"🛡️","tokens_out":6918,"duration_ms":69715,"temperature":0.7,"pith_summary":"The paper introduces WeClawArena, a benchmark and runtime sandbox for situations in which several human owners each delegate an AI agent, and those agents must complete one collaborative tool-use task while each sees only its own private workspace. The central claim is that task utility and attack harm are separate outcomes: an agent team can finish the task while leaking private information, trusting poisoned evidence, or accepting an invalid approval, and it can also fail the task without any attack having caused final harm. To test this, WeClawArena contributes 124 base tasks across six domains, expanded into 620 matched scenarios with a benign control and four attack-vector variants each. The sandbox records messages, tool calls, resource operations, governance decisions, and final workspace states, and the evaluation reports deterministic task success and LLM-judged attack success separately. The empirical split motivates the design: of 3,743 judged attack rows, 1,152 reached attack success, but only 117 of those also completed the task.","feed_headline":"AI agents can finish the job and still leak data","feed_subtitle":"A 620-scenario benchmark measures task utility and attack harm separately for multi-user agent networks.","key_machinery":"The load-bearing mechanism is the division of labor between a deterministic verifier and an offline attack judge. The verifier scores task utility from final workspace states and task contracts; the judge scores attack success from a bounded evidence packet assembled after each run, containing scenario metadata, attack metadata, transcript messages, tool calls and observations, task-score fields, and owner or governance context. An attack overlay may add messages, files, or database rows, but cannot directly edit the final score state, so any judged harm must have traveled through ordinary agent actions. This separation is what allows the paper to conclude that harm and utility are independe","core_discovery":"The paper's central claim is that collaborative tool use among agents who each serve a different human owner creates a distinct failure surface: a team can complete the visible task while leaking a private budget, trusting poisoned evidence, or accepting an invalid approval. WeClawArena operationalizes this claim as a benchmark and sandbox, with 124 base tasks across bargaining, bidding, travel, software-engineering workspaces, clinical coordination, and trading, each expanded into one benign control and four attack-vector variants, for 620 scenarios. Utility is scored deterministically from final workspace states and contract predicates, while attack success is judged offline from bounded e","pith_inferences":["I would expect the strongest practical leverage to be structural: because invalid-authority paths were the largest judged harm category, an agent-network gateway could enforce approval and ownership checks at the tool layer rather than relying on the model to refuse; the governance rows give a direct test of such enforcement.","The evidence-packet design generalizes beyond attacks: the same bounded records could support counterfactual auditing (what would have happened without the attack) and post-hoc explanation of ordinary task failures, which the paper leaves implicit.","A cheap deployment test suggested by the data: monitor not just final task outcomes but evidence links that indicate poisoned inputs or authority-path breaks, since many successful attacks preserved utility.","The gap between frontier and open-weight models in judged attack resistance points to attack pressure as a training axis rather than purely a prompting-time property; this is my inference, not the paper's claim."],"forward_implications":["Benchmarks for multi-agent collaboration should report utility and attack success separately; a utility-drop number is a symptom, not a substitute for an audit of final harm.","Attack harm is not one thing: in the sweep, governance and security attacks succeed far more often than collaboration attacks, and the same harm surface lands very differently across domains, so safety evaluation should be per-surface rather than a single scalar.","The 117 runs that both completed the task and reached attack success are deployment-relevant: completion signals alone would miss privacy leakage, poisoned evidence, and invalid authority paths.","Judged attack-resistance rank can be stable across judge choices, with the overall model ordering unchanged when a second judge re-reads the same evidence packets, which supports bounded-evidence ASR as a reproducible audit layer.","The benign-control false-positive rate stays below one percent for both judges, meaning the audit layer could serve as a low-noise filtering signal, not only as a headline metric."],"supporting_citations":[{"why":"Supplies the persistent personal-agent harness used as the runtime substrate for owned-agent behavior in the simulations.","marker":"(Steinberger and OpenClaw Contributors, 2025)"},{"why":"Supplies the human-centered agent-network framing and privacy-risk motivation that WeClawArena shifts from social conversation to verifiable tool-use collaboration.","marker":"(Wang and Jiang, 2026a)"},{"why":"Defines the tool-agent-user evaluation paradigm that WeClawArena extends from one user to multiple owner-scoped workspaces.","marker":"(Yao et al., 2025)"},{"why":"Extends tool-agent benchmarking to dual-control environments, providing the closest prior scoring setting the paper adapts.","marker":"(Barres et al., 2025)"},{"why":"Provides contextual integrity, the standard that defines when an information flow counts as privacy harm in the attack-success rubric.","marker":"(Nissenbaum, 2004)"},{"why":"Supplies the LLM-as-a-judge methodology whose reliability constraints motivate bounded evidence packets and offline judging.","marker":"(Zheng et al., 2023)"},{"why":"One of the privacy-in-agent-memory benchmarks whose harm surfaces motivate the privacy attack vector.","marker":"(Juneja et al., 2025)"},{"why":"Supplies the auditable-agent evidence-recording perspective that motivates the runtime evidence-packet design.","marker":"(Nian et al., 2026)"}],"fun_headline_variants":["Agents finish tasks but leak secrets in 620 scenarios","AI agent teams can do the job and still leak data","Benchmark: agents complete tasks, but at a cost","620 scenarios show agent teams can work and harm","Multi-agent collabs: job done, secrets out"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The headline attack-success numbers rest on a single LLM judge reading evidence packets that contain the attack's 'expected exposure path' and 'evidence hints'; if the judge is tracking those hints rather than independently establishing final harm, the ASR values and the model-resistance ranking could shift.","fun_headline_variants_meta":{"raw":{"variants":["Agents finish tasks but leak secrets in 620 scenarios","AI agent teams can do the job and still leak data","Benchmark: agents complete tasks, but at a cost","620 scenarios show agent teams can work and harm","Multi-agent collabs: job done, secrets out"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1160,"prompt_tokens":770,"completion_tokens":390,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":312}},"tokens_in":514,"tokens_out":390,"duration_ms":4435,"temperature":1.0,"reasoning_tokens":312,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:51:28.882616+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An audit study that removes the 'expected exposure path' and 'evidence hints' from the evidence packets and has a judge or blinded human annotators re-label the 200-row pilot: if ASR drops substantially or agreement collapses, the original judge was leaning on the attack description rather than on recorded evidence of final harm.","supporting_citations":[{"cited_title":"2025 , howpublished=","cited_arxiv_id":null,"evidence_quote":"Supplies the persistent personal-agent harness used as the runtime substrate for owned-agent behavior in the simulations."},{"cited_title":"2025 , note=","cited_arxiv_id":null,"evidence_quote":"Defines the tool-agent-user evaluation paradigm that WeClawArena extends from one user to multiple owner-scoped workspaces."}],"review_version":1}