{"id":"5c468785-846c-43dd-b005-7b12bfa51dcc","arxiv_id":"2607.18826","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Cross-agent asynchronous attack sessions can be linked at 0.82 pairwise AUC from proxy-visible tool-use and prompt-style residue in the authors' synthetic SCD-v1 benchmark, far above adapted per-session detectors and chunked LLM judges.","lead":"This paper proposes that attacks spread across many AI agents over time can be linked by a shared security proxy that watches tool use and writing style, and builds a controlled benchmark to test it. On that benchmark, a simple fingerprint score separates same-campaign attack pairs from unrelated sessions at 0.82 AUC, while per-session detectors and chunked LLM judges stay near chance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Primary AUC excludes benign traffic, so the 'benign-heavy' qualifier in the central claim is not directly measured; all-pair diagnostics show a much weaker operating point.","rationale":"The reader's conditional verdict is well supported. I looked for the place where the central claim's conditions and the evidence diverge. The contribution is explicitly about 'controlled, benign-heavy, asynchronous conditions,' but the primary pairwise-AUC protocol defined in §4.1 uses only campaign-vs-campaign negatives; benign and isolated sessions never enter the denominator. The all-session V-measure (0.269) and all-pair Precision@20/Recall@100 (2/20, 5/1728) are the only benign-heavy readouts and are much weaker than 0.818. This is a claim/evidence alignment issue that can be settled by a single computation on the released artifacts; it does not depend on real-world transfer assumptions. External validity (the reader's weakest assumption) remains a legitimate secondary concern, but the metric gap is more load-bearing because it affects whether the controlled benchmark itself demonstrates the stated claim. The disclosed E10 selection and style-crossing drop are real but secondary; they are reported transparently and do not by themselves change the verdict. I would not reject: the task formalization, leakage audits, and fixed-weight reproducibility are genuine contributions. I would keep the paper CONDITIONAL and require the all-pair AUC (or a revised claim) before treating the benign-heavy clause as established.","tokens_in":22963,"tokens_out":7846,"duration_ms":69824,"concrete_test":"On the released Full 2000 artifact, recompute pairwise AUC using the fixed w*=(0.6,0,0.4) scores with positives = same-campaign pairs and negatives = all other pairs (benign–benign, benign–isolated, benign–campaign, isolated–isolated, isolated–campaign, cross-campaign). Compare this all-pair AUC to the restricted campaign-only 0.818; also report Precision@10 and Recall@100 on the same full pair set. If all-pair AUC drops below ~0.7, the 'benign-heavy nontrivial' clause of the central claim is not supported by the primary protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is scoped to 'controlled, benign-heavy, asynchronous conditions' (Intro, Contributions), and the abstract's headline is 0.82 pairwise AUC. But the scoring protocol in §4.1 defines a negative pair as 'two campaign sessions from different campaigns,' explicitly excluding benign sessions and isolated attacks. The 0.818 number is therefore a campaign-vs-campaign ranking, not a benign-heavy ranking. The only benign-heavy readouts are the all-session V-measure (0.269) and the all-pair alert-budget diagnostics (Precision@20 = 2/20, Recall@100 = 5/1,728 over ~2M candidate pairs, Table I). No threshold-free all-pair AUC with benign and isolated negatives is reported. The paper is honest that these are diagnostics, but the contribution sentence promises nontriviality in a benign-heavy regime; the primary metric does not test that regime. If the all-pair AUC is materially below 0.818 (e.g., 0.6–0.7), the claim overstates what SCD-v1 demonstrates, and the 'distinct layer' would need to be repositioned as a post-filter correlation aid, which is a weaker claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a new security-evaluation task: cross-agent asynchronous campaign attribution, in which sessions from the same latent adversarial campaign are linked from proxy-observable prompt, tool-call, and timing residue, without shared runtime state, campaign labels, or attacker identity. The authors define a five-condition threat model, introduce A2FV as a fixed, transparent pairwise scoring protocol over structural, temporal, and stylometric residue blocks, and construct SCD-v1, a synthetic persona-matched corpus with 70% benign traffic, isolated attacks, and multi-session campaigns. On SCD-v1 they report 0.818 pairwise campaign-linking AUC, while score-delta adaptations of per-session detectors (0.51–0.52) and chunked LLM judges (0.50–0.51) remain near chance. The paper also reports matched-persona L0/L1/L2 evasion sweeps, leakage audits, Qwen/Gemini regenerations, crossed-style hard-slice controls, and OpenClaw/LangGraph native-framework probes, and it is unusually explicit about limitations.","tokens_in":23125,"tokens_out":7557,"duration_ms":68503,"significance":"If the 0.818 AUC holds, the paper's main contribution is significant: it defines a previously missing evaluation layer and provides a reproducible benchmark and reference protocol. The strengths are the layered evidence: bootstrapped confidence intervals, leakage audits, matched-persona L0/L1/L2 triplets, crossed-style controls, independent-generator regenerations, native-framework probes, and released artifacts. The unweighted structural+prompt linker (0.825) independently supports the existence of proxy-observable campaign signal. However, the headline number must be read with three restrictions: the primary AUC excludes benign negatives; the three A2FV weights are selected on the same SCD-v1 corpus; and the crossed-style hard slice drops the fixed score to 0.599. These restrictions do not invalidate the contribution, but they require the claims to be re-scoped.","major_comments":[{"comment":"The headline pairwise AUC (0.818) uses negatives defined as campaign sessions from different campaigns; benign sessions and isolated attacks are excluded unless an experiment explicitly includes them. The contribution in §1 nonetheless claims nontriviality “under controlled, benign-heavy ... conditions.” The only benign-heavy readouts are all-session V-measure 0.269 and all-pair Precision@20=2/20, Recall@100=5/1,728 over ~1.999M pairs (Table I); no threshold-free all-pair AUC with benign and isolated negatives is reported. If that all-pair AUC is materially below 0.818, the abstract’s “0.82” overstates the benign-heavy claim. Please report all-pair AUC with benign/isolated negatives or re-scope the claim to campaign-vs-campaign ranking.","section":"§4.1 Metrics; Table 2; §1 Contributions"},{"comment":"The weights w* are selected on the SCD-v1 calibration artifact by maximizing NMI against campaign labels at 0.1 grid resolution, and the headline 0.818 is reported on the Full 2000 corpus. Unless the calibration artifact is a strictly disjoint split, this is in-sample weight selection. The paper’s own unweighted structural+prompt control (0.825) shows the signal is not an artifact of tuning, so this is not fatal; but the calibrated A2FV number should be presented as a same-corpus selected-weight result, with a disjoint-split or fixed-default-weight row as the primary deployment-style estimate.","section":"§3.3 Eq. (3); §4.2 Table 2"},{"comment":"In SCD-v1, campaign identity is generated jointly with a persona style directive; positives are same-persona/campaign pairs and negatives are different-persona pairs. The distributional audit shows style attributes are recoverable (style AUC 0.928), and the fresh crossed-style hard slice drops fixed A2FV to AUC 0.599, stylometry-only to 0.271, and structural-only to 0.686. The 0.818 headline is therefore substantially style-correlated. The paper does state this, but the abstract still presents 0.82 as the headline. I recommend reporting the crossed-style hard-slice and structural-only hard-slice alongside the headline, or adjusting the abstract so that the style-sensitive nature of the 0.82 number is not buried.","section":"§4.1 matched-persona design; Finding 4, Table J"},{"comment":"The limitation paragraph correctly concedes that adversaries who deliberately imitate normal operator workflows would remove the residue channel on which A2FV relies, and that real telemetry degradation (missing tool events, coarse timestamps, taxonomy shifts) is unmeasured. I do not regard this as an internal inconsistency, but it should be reflected in the conclusion’s wording: the paper establishes a controlled-benchmark phenomenon, not a deployment-ready layer. The current conclusion says “establishes the missing evaluation layer,” which is acceptable if read as the protocol/benchmark, but the abstract’s “in the wild” phrasing overreaches.","section":"§5 Scope and limitations"}],"minor_comments":[{"comment":"Rows use per-K best w* at grid resolution 0.1; the caption should state explicitly that these are not fixed-weight results, so readers do not read the sweep as evidence of fixed-weight robustness.","section":"Table E"},{"comment":"The “Unweighted structural+prompt control” row is a strength, but the text should make explicit that this row, rather than the calibrated A2FV row, is the cleanest evidence that the task exposes proxy-observable signal, because it has no fitted weights.","section":"§4.2, Table 2"},{"comment":"The text says “Raw Precision@20/Recall@100 ... are harsh all-candidate pre-filter diagnostics”; please state candidate-pair and positive counts in the main text, not only in Table I, so the 2/20 and 5/1,728 numbers are interpretable without hunting for the appendix.","section":"§4.1 Metrics"},{"comment":"The notation is inconsistent: the abstract and some sections use $A^2FV$, while the body uses A2FV. Please unify.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually honest and well-supplied with controls. The main risk is over-claiming “benign-heavy” and “distinct layer” given that the primary pairwise metric excludes benign traffic and the weights are selected on the same corpus. The unweighted control and crossed-style hard slice should be used to recalibrate the abstract. I would be happy to see a revised version that adds the all-pair AUC, clarifies the calibration split, and reframes the headline."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the two things worth knowing. This paper defines a genuinely new task—linking asynchronous attacks across independent LLM agents from proxy-observable residue only—and backs that formalization with an unusually careful, self-critical evaluation. The headline 0.82 pairwise AUC is a campaign-vs-campaign number, not a measure of how A2FV performs in benign-heavy traffic. The stress-test note is right: the primary metric excludes benign sessions and isolated attacks, so the abstract's 'benign-heavy' phrase overstates what the main result actually demonstrates. The all-session V-measure (0.269) and all-pair precision/recall diagnostics (2/20 at precision@20) are in the paper but are much weaker operating points.\n\nWhat is genuinely new: Definition 1's five-condition threat model is a real gap in the literature, and Table 1 backs that up. SCD-v1 is a controlled, persona-matched benchmark with layered controls: leakage audits, matched L0/L1/L2 triplets, crossed-style diagnostics, independent Qwen/Gemini regenerations, and native-framework probes. The finding that per-session detectors and chunked LLM judges stay near chance under the same pairwise protocol is supported and useful. The authors are also unusually honest about scope—they repeatedly say SCD-v1 is synthetic, that the signal is partly style-correlated, and that fully adaptive L3 is out of scope. That honesty is earned.\n\nThe soft spots, in proportion. First, the in-sample weight selection: the three block weights are chosen on SCD-v1 by maximizing NMI against the campaign labels, and the reported 0.818 is on the same corpus. The unweighted structural+prompt control at 0.825 is reassuring, but the headline is still a fitted number. Second, the crossed-style hard slice drops the fixed score to 0.599; structural residue survives at 0.686, so the signal isn't pure style, but the abstract's confidence is higher than the controlled result supports. Third, Appendix I discloses a stronger internal E10 variant at 0.680 that was not published—not a scandal, but a reminder that the released stress result is the smaller drop. Fourth, and most structurally, the weakest assumption is external validity: SCD-v1's synthetic traffic must resemble real adversarial and benign statistics, and the paper's own Section 5 concedes that an adversary imitating normal operator workflows removes exactly the residue A2FV relies on. So the contribution is a reproducible reference protocol and a benchmark, not a deployed correlate.\n\nWho should read this: anyone working on agent security evaluation, proxy-layer monitoring, or incident triage. It deserves serious refereeing—the task definition will likely be reused, and the evidence, though synthetic, is internally consistent and honestly reported. I would send it to peer review with a request for an out-of-sample regenerated test set and a threshold-free all-pair AUC that includes benign and isolated negatives. My own verdict is close to the reader's: conditional. The core task definition is valuable; the headline claim needs to be repositioned to match what the primary metric actually measures.","headline":"A genuinely new task formalization and an unusually honest benchmark, but the headline AUC is a campaign-only number and the weights are fit on the same corpus.","tokens_in":23776,"tokens_out":3436,"would_cite":true,"duration_ms":29242,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Proxy-observable tool-use, timing, and prompt residue suffice to link asynchronous LLM-agent sessions from one latent campaign, reaching 0.82 pairwise AUC on a controlled benchmark while per-session score linkers stay near chance.","keywords":["cross-agent campaign attribution","LLM agents","prompt injection","proxy observability","pairwise AUC","tool-use fingerprinting","stylometry","asynchronous attacks"],"falsifier":"A deployment trace with independently established campaign identity (for example, incident-response ground truth) in which the adversary deliberately reuses normal operator tool sequences and prompt style: if A2FV's pairwise AUC on such pairs falls to near chance while per-session detection stays strong, the residue channel is insufficient in the wild. More cheaply, on a crossed-style slice larger than 120 sessions, if the structural-only hard-slice AUC drops from about 0.686 toward 0.5 as persona variance grows, the structural lower-bound signal fails.","tokens_in":22680,"feed_emoji":"🔗","tokens_out":4561,"duration_ms":38502,"temperature":0.7,"pith_summary":"The paper claims that defending LLM agents session-by-session misses a real class of attack: one adversary distributing related injections across independent agents and teams, asynchronously, so each local guardrail sees only a fragment. It formalizes this as cross-agent asynchronous campaign attribution — a linking task, not a per-session classification task — and shows that a lightweight proxy-side fingerprint protocol, A2FV, can rank same-campaign session pairs above unrelated pairs using only proxy-observable residue: tool-use structure, inter-call timing, and prompt surface style. On the paper's constructed persona-matched benchmark SCD-v1, A2FV reaches pairwise AUC 0.818, while adapted per-session detectors and chunked LLM judges stay near chance. The contribution is the task definition plus a transparent reference protocol and benchmark, not a worst-case robustness guarantee; the paper is explicit that fully adaptive adversaries with detector-score feedback remain out of scope. A sympathetic reader would care because this defines and measures a missing evaluation layer for agent security.","feed_headline":"Same-campaign LLM-agent attacks linked at 0.82 AUC","feed_subtitle":"Tool-use, timing, and prompt residue let a shared proxy spot one adversary across many sessions.","key_machinery":"Asynchronous Attribution Fingerprint Vectors (A2FV): a transparent, label-free-at-inference scoring protocol. Each session is embedded in three blocks — structural (hashed counts of decorated tool-call n-grams for n=1,2,3, failure-pivot transitions, and a tactic histogram), temporal (tool count, duration, log-gap moments, dip statistic, IQR/sigma, payload-size moments), and stylometric (TF-IDF over character and byte n-grams of the prompt) — normalized and compared by cosine similarity per block, then combined by three scalar weights selected once on a calibration split; the fixed deployment weights are (0.6 structural, 0.0 temporal, 0.4 stylometric). The work it does is to produce a pairwis","core_discovery":"The central discovery is that the identity of a latent adversarial campaign leaves recoverable residue in proxy-observable logs even when sessions share no runtime state, arrive asynchronously, and are interleaved with benign traffic. Defining each session by three channels — structural residue (decorated tool-call n-grams, failure-pivot transitions, tactic histogram), temporal residue (inter-call gap moments and payload-size statistics), and stylometric residue (character and byte n-gram TF-IDF over the prompt) — the paper's A2FV protocol forms a weighted cosine score and shows that linking by this score separates same-campaign pairs from different-campaign pairs (AUC 0.818, 95% CI [0.797,","pith_inferences":["An implication the paper leaves implicit: if the residual-channel hypothesis generalizes, the same pairwise-linking protocol could be applied to other sparse telemetry such as API-login patterns or cloud-console actions, with A2FV-style structural and stylometric blocks.","The measured style-sensitivity (hard-slice drop from 0.818 to 0.599, stylometry-only collapse to 0.271) suggests that real-deployment AUC will depend heavily on how much benign traffic shares the adversary's surface style; a style-normalization pre-layer might recover part of that gap.","Because temporal residue is weak in this benchmark yet the threat model includes timing, the paper's own logic implies richer traces with queueing, retries, and fallback delays could provide the strongest new signal; re-running weight selection on such traces is a testable extension.","The deepest threat implicit in the limitations section is an adversary who deliberately imitates normal operator workflows; the residual-channel hypothesis predicts such mimicry would attack structural and stylometric blocks simultaneously, so a benchmark with a mimicry persona is a natural next stress test."],"forward_implications":["Campaign attribution is a measurable layer distinct from per-session safety classification, so agent-security evaluations should include pairwise and partition linking protocols, not only per-session scores.","Score-only adaptations of per-session detectors discard most identity residue: the best score-delta linker reaches 0.522 AUC versus 0.818 for A2FV, so local safety outputs cannot be directly repurposed as campaign linkers.","Chunked LLM judges fail at scale (0.50–0.51 AUC) because global campaign identity fragments across chunk boundaries; a cheap fingerprint first pass followed by LLM or analyst review on ranked clusters is the intended operational pattern.","Static non-oracle evasion (L1/L2) leaves pairwise AUC essentially intact (0.807–0.816), and a dimension-aware non-oracle composition keeps fixed-weight AUC at 0.833, suggesting the residue bundle is not removed by one or even three simultaneous static transforms.","Fixed-weight signal transfers to independent generator regenerations (0.792/0.826 AUC) and to controlled native-framework traces (0.66–0.71 AUC) after schema mapping, making a pure benchmark-artifact explanation less plausible."],"fun_headline_variants":["Same-campaign LLM sessions linked via proxy logs at 0.82 AUC","Proxy-side fingerprint links asynchronous LLM-agent attacks","A2FV: spotting one campaign across many LLM sessions","Tool-use and prompt residue expose hidden LLM campaigns","Link cross-agent attacks without shared state"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The measured 0.818 AUC transfers from SCD-v1's controlled, persona-matched synthetic generation to real adversarial campaigns and benign operator traffic; the paper itself concedes that real telemetry could degrade A2FV through missing tool events, coarse timestamps, changing tool taxonomies, and adversaries who deliberately imitate normal operator workflows.","fun_headline_variants_meta":{"raw":{"variants":["Same-campaign LLM sessions linked via proxy logs at 0.82 AUC","Proxy-side fingerprint links asynchronous LLM-agent attacks","A2FV: spotting one campaign across many LLM sessions","Tool-use and prompt residue expose hidden LLM campaigns","Link cross-agent attacks without shared state"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00046,"raw_usage":{"total_tokens":2165,"prompt_tokens":793,"completion_tokens":1372,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":1292}},"tokens_in":537,"tokens_out":1372,"duration_ms":23396,"temperature":1.0,"reasoning_tokens":1292,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T14:10:49.629405+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A deployment trace with independently established campaign identity (for example, incident-response ground truth) in which the adversary deliberately reuses normal operator tool sequences and prompt style: if A2FV's pairwise AUC on such pairs falls to near chance while per-session detection stays strong, the residue channel is insufficient in the wild. More cheaply, on a crossed-style slice larger than 120 sessions, if the structural-only hard-slice AUC drops from about 0.686 toward 0.5 as persona variance grows, the structural lower-bound signal fails.","supporting_citations":[],"review_version":1}