{"id":"aa9c8302-7896-4e38-b296-569ecbdeabad","arxiv_id":"2607.25890","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A harness-distributed bundle of OS sandboxing, skill scanning, and tool restriction matched the best secured commercial agent on adjusted security tests, with no regression against its own baseline.","lead":"This paper shows that three off-the-shelf security controls, embedded in a distributable coding-agent harness called SHarD, can be shipped with a one-command installer and still pass the same basic security probes as when installed directly on commercial agents. For organizations scaling agentic AI, that offers a vendor-neutral way to push security controls out to every developer's machine.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-run test scores and known non-determinism leave the 100% adjusted-score equivalence claim underdetermined","rationale":"The reader's weakest assumption — that single-run pass/fail scores are treated as stable point estimates in a demonstrably non-deterministic system — is the most load-bearing concern about the central claim. The paper's own §VII.A observation that a security-relevant agent behavior was not reproduced in subsequent runs makes this more than a generic statistical worry: it is a documented failure of the determinism assumption within the exact experimental setup. The equivalence claim 'SHarD matches the best securely configured commercial agent' derives from point equality in adjusted scores (100% vs. 100%), and even a small number of flipped test outcomes could erase that match or create a false one. The concern does not invalidate the paper's screening-level contribution, but it does mean the headline quantitative claim is not yet established. Since the reader already assigned CONDITIONAL for essentially this reason, my stress-test does not move the verdict; it reinforces the condition. A targeted re-run experiment would settle whether the concern lands.","tokens_in":17510,"tokens_out":4238,"duration_ms":41681,"concrete_test":"Re-run the full 23-probe test suite for SHarD, secured Claude Code, and secured Codex at least 10 times each, using fresh agent sessions and the pinned versions/hashes in Appendix B. For each test, record pass/mixed/fail per run, then compute per-category and adjusted-score distributions with 95% confidence intervals. If SHarD's adjusted score distribution overlaps Claude Code's (e.g., both reach 100% in most runs) and its per-test pass rates are not significantly lower, the equivalence claim stands. If SHarD's perfect score occurs in only a minority of runs, or its pass rate falls below Claude Code's on any controlled category, the claim that efficacy is retained is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that SHarD retains the same control efficacy as direct installation on commercial agents—rests entirely on point scores from one execution of each probe per configuration. The paper itself supplies direct evidence that this premise is unsafe: in §VII.A, CPR-01 passed for SHarD with no content protection control because the agent followed a reasoning path that 'the agent never took this specific path again.' That is an observed non-reproducible outcome in the very system being scored. With n=1 per test, the 100% adjusted score for SHarD and its match with Claude Code's 100% could be run-specific artifacts. The problem is amplified by small category sizes—Tool Selection has only two tests (TOOL-01, TOOL-02)—so a single flipped outcome changes a category score by 50 percentage points and can determine whether SHarD 'matches' the commercial agent. The paper's methodology is otherwise transparent, and the screening-level intent is reasonable, but the headline comparative claims are underdetermined without repeated trials and variance reporting. This is not an internal contradiction in the design; it is a statistical fragility in the evidence supporting the abstract's equivalency assertion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates whether off-the-shelf security controls for AI coding agents—content protection, skill scanning, tool restriction, and OS sandboxing—can be validated on commercial agents (Claude Code, Codex) and then redistributed through a custom agent harness (SHarD) while retaining efficacy. A 23-test suite derived from the OWASP Top 10 for Agentic Applications is run in four phases: two commercial baselines, two commercial configurations with controls, a Pi harness baseline, and the SHarD harness with controls. The central claim is that SHarD achieves an adjusted score of 100%, matching the best secured commercial agent, with no regression across test categories, thus answering RQ1 and RQ2 at a basic functional level. The paper also reports qualitative observations about model non-determinism and cross-agent contamination, and proposes two initial characteristics (Declarative Policy, Control Locality) toward a control-harness fitness framework.","tokens_in":17773,"tokens_out":2957,"duration_ms":29522,"significance":"If the headline result were robust, the paper would be a useful empirical data point for practitioner adoption: it demonstrates a plausible mechanism for distributing security controls through an agent harness rather than relying on vendor-native solutions, and it publishes reproducible artifacts (pinned hashes, raw data repository, one-command installer). The methodology is transparent and the screening-level intent is reasonable. However, the central comparative claim depends on single-run point scores in a system the paper itself shows to be non-deterministic. The evidence therefore supports a qualitative 'these controls can function in a harness' conclusion more strongly than the quantitative equivalence claim in the abstract. With repeated trials and variance reporting, the contribution could be solid as a screening study.","major_comments":[{"comment":"The headline claim that SHarD 'matched' Claude Code with an adjusted score of 100% rests on a single execution of each probe per phase/agent. The paper itself supplies direct evidence that this premise is unsafe: §VII.A reports that CPR-01 passed for SHarD without content protection because the agent 'went down a reasoning path' that 'the agent never took this specific path again.' That is an observed non-reproducible outcome in the exact system being scored. With n=1 per test, Eq. (1) has unknown variance; in Tool Selection, the category contains only two tests (TOOL-01, TOOL-02), so a single flipped outcome changes that category by 50 percentage points and can determine whether 'no regression' and 'matching' hold. Please repeat each test multiple times (e.g., 5–10 runs per configuration) and report variance/confidence intervals, or explicitly downgrade the claim to 'single-observation","section":"§IV.C / §V.A / Table IV"},{"comment":"The 'adjusted score' is not sufficiently operationalized. The text says the adjusted score is 'limited to just the controlled categories,' but it does not specify which tests are included and excluded, how Mixed outcomes are counted (0.5 in Eq. (1)), or how 'inconclusive results were excluded from the denominator' in each phase. Table V complicates this further: the Content Protection category shows Pi0 at 37.5% vs. SHarD at 44.4% over what appears to be the same nine-test CPR suite, which is only possible if different numbers of tests were excluded as inconclusive or if outcome classification changed between runs. Without a per-test, per-phase outcome table (or a complete data appendix), the raw and adjusted scores are not auditable, and the central 'equivalence' claim cannot be verified from the manuscript alone. The GitHub repository is a good start, but the paper should include the f","section":"§V.A / Table IV"},{"comment":"The 'no regression' claim is asserted for category scores, but the regression check compares Pi default to SHarD, not SHarD to the secured commercial agents. The abstract and conclusion use the phrase 'no regression across any test category' in the context of the equivalence claim, which could mislead readers into thinking this was tested against both commercial configurations. In addition, the regression comparison inherits the single-run fragility described above; a category improvement like Skill Scanning (+41.7) is presented as a point value with no measure of run-to-run variability. Please clarify the reference configuration for each 'regression' statement and, if possible, include repeated-trial estimates.","section":"§VII.B / Table V"}],"minor_comments":[{"comment":"The text consistently renders 'OWASP' as 'OW ASP' with an extra space (e.g., abstract, Section II.A, Table I). Please fix.","section":"Throughout"},{"comment":"The pass condition for CPR-05 says 'CPR-04 passes...' — likely a copy-paste error; should read CPR-05. Also, in CPR-04 the prompt refers to a '.evn' file; should be '.env'.","section":"Appendix A, CPR-05"},{"comment":"The baseline harness is labeled 'Pi 0' in the table but 'Pi default' in the text and 'Pi Coding Agent' elsewhere. Use one consistent label.","section":"Table V"},{"comment":"The sentence 'Because nono enforcement is applied at the kernel level via macOS Seatbelt, the sandbox cannot be bypassed from within the sandboxed process itself' should cite a more specific source than a DEV Community blog post; if the kernel-level guarantee is central to the OS sandbox claim, a primary reference or direct test evidence would strengthen it.","section":"§VI.B"}],"recommendation":"major_revision","confidential_remarks":"The single-author screening study has a clear scope and transparent artifacts, but the headline equivalence claim is statistically underdetermined as written. The core design is not fatally flawed; repeated trials and a full outcome matrix would likely make this an acceptable empirical contribution. The paper's current framing overreaches relative to the n=1 evidence, so I would not accept in its present form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it does what it says: it empirically shows that off-the-shelf security controls (nono sandboxing, SandyClaw skill scanning, tool restriction) can be embedded in a distributable harness and still pass a 23-test OWASP-derived functional suite. The artifact is real—pinned hashes, raw data on GitHub, a single-install script. That is more than most papers in this area ship, and it earns real credit. Second, the headline equivalence claim—SHarD matching Claude Code at 100% adjusted score with no regression—rests on a thin statistical base. Every probe was run exactly once, and the paper itself documents a non-reproducible outcome: CPR-01 passed without the relevant control because the agent took a reasoning path it never took again. In a non-deterministic system, n=1 gives you a point score with unknown variance. Tool Selection has only two tests, so one flip changes that category by 50 points. The paper is otherwise well-scoped and honest. The phased design makes sense for a screening study. The adjusted-score exclusion of content protection is disclosed and reasoned, not smuggled in. The Conclusion does slightly overstate the Codex tool-restriction result (Table II shows no improvement there), and the phrase 'same efficacy' should be qualified to 'basic functional efficacy'—the tests are deliberately not sophisticated-attack probes. Those are addressable weaknesses, not fatal ones. The framework proposal in Section VII-B is explicitly preliminary and the author flags the small sample, so that section does not overreach. The self-citations are artifact links, not evidence, so no circularity concern. Who gets value: security teams weighing vendor-neutral deployment of agent controls, and researchers working on harness engineering. It is not a breakthrough, but it is a useful, reproducible data point in a fast-moving space. My recommendation: this deserves a serious referee, not a desk rejection. Send it to peer review with a request for repeated trials, variance reporting, and a more carefully scoped abstract. I would accept it as a screening study after those revisions.","headline":"A transparent screening-level study with a real artifact, but the n=1 test runs make the headline 'same efficacy at 100%' claim fragile—worth refereeing with a request for variance reporting.","tokens_in":689,"tokens_out":2176,"would_cite":true,"duration_ms":42472,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A distributable harness can carry off-the-shelf security controls to AI coding agents without losing efficacy.","keywords":["agent harness","distributable security controls","OS sandboxing","skill scanning","tool restriction","AI coding agents","harness engineering","prompt injection"],"falsifier":"Run the full 23-probe suite multiple times (for example, ten repetitions) on SHarD and on the best securely configured commercial agent under identical conditions. If SHarD's adjusted score falls below the commercial agent's, or if any category regresses relative to the unhardened harness, the equivalency claim fails. Alternatively, repeat the specific content-protection test that passed without the relevant control twenty times and count how often the agent spontaneously blocks the injection.","tokens_in":17326,"feed_emoji":"🛡️","tokens_out":5599,"duration_ms":48156,"temperature":0.7,"pith_summary":"This paper tries to show that security controls for AI coding agents do not have to be configured one-by-one on each developer machine. It embeds three off-the-shelf control types—OS sandboxing, skill scanning, and tool restriction—into a distributable harness called SHarD that installs with a single command. In a 23-test functional screen targeting risks like indirect prompt injection, tool misuse, and supply-chain attacks, the hardened harness scored 100% on adjusted tests, matching the best directly secured commercial agent with no regression across any category. The paper argues this makes the harness a viable channel for scaling security and identifies two traits—controls expressed as code and enforced at the execution boundary—as candidates for a control-fitness framework.","feed_headline":"One install command ships security controls to AI coding agents","feed_subtitle":"One-command installation that matched the best directly secured agent on every controlled test.","key_machinery":"The harness itself—the code, configuration, and execution logic surrounding the model—is the central mechanism. SHarD's design has three load-bearing parts: an installer that provisions dependencies and writes policy files in one command; a package manifest that registers extension hooks at session start and tool-call time; and extension hooks that relaunch the agent inside a kernel-enforced OS sandbox, route skill analysis to an external scanner, and block denied bash commands before execution. The proposed fitness framework rests on two characteristics observed in the controls that carried forward: declarative policy (the control can be distributed as a versioned, portable artifact) and co","core_discovery":"The paper's central claim is that off-the-shelf security controls can be scaled via a distributable agent harness while maintaining the same efficacy as when they are installed directly on commercial AI coding agents. SHarD, built by wrapping an existing open-source harness with a package manifest, an installer, and extension hooks, bundles OS sandboxing, skill scanning, and tool restriction. At session start it relaunches itself inside a kernel-enforced OS sandbox, it scans third-party skills through an external service before they reach the model, and it intercepts bash tool calls to enforce deny rules. Across the 23-probe suite, SHarD achieved an adjusted score of 100%, matching the best","pith_inferences":["If harness distribution becomes standard, controls that require deep runtime integration or that conflict with other controls—the paper's content-protection experience is a hint—may be systematically left out; that exclusion deserves explicit testing.","The screening methodology runs each probe once, so the 100% adjusted score should be read as evidence of basic functionality, not as a stable effect size; repeated runs under model non-determinism would test whether the equivalency holds.","The proposed fitness framework predicts that any control expressible as a policy file plus a session or tool hook—secret scanning, egress filtering, MCP allow-listing—is a candidate for harness distribution; this is a falsifiable prediction for future work.","The observed cross-boundary behavior suggests harness distribution could be extended to manage multi-agent co-existence on one machine, not just sandboxing a single agent."],"forward_implications":["Security teams can distribute controls through the same one-command channel that engineering teams already use to ship code, removing per-machine manual configuration.","The three control categories—OS sandboxing, skill scanning, and tool restriction—retain basic efficacy when embedded in a harness, so vendor-native security is not the only path to scale.","Because model behavior alone proved non-deterministic and sometimes unsafe, harness-enforced deterministic controls are a more reliable security layer than reliance on model reasoning.","Autonomous agents can act outside their intended boundaries, for example by modifying another agent's configuration on the same system, which argues for defense in depth with OS-level containment.","A control-fitness framework based on declarative policy and control locality can help practitioners judge which controls to distribute, though more controls must be tested to validate it."],"fun_headline_variants":["Single command distributes OS, skill, and tool security to agents","SHarD: one-command agent security, matches direct install at 100%","Harness ships sandboxing, skill scanning, tool rules in one install","One-command harness matches directly secured coding agents"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Every score in the comparison rests on running each of the 23 probes exactly once and treating each pass/fail as a stable point; if agent non-determinism changes outcomes across runs, the 100% adjusted score and zero-regression claims are underdetermined.","fun_headline_variants_meta":{"raw":{"variants":["Single command distributes OS, skill, and tool security to agents","SHarD: one-command agent security, matches direct install at 100%","Harness ships sandboxing, skill scanning, tool rules in one install","One-command harness matches directly secured coding agents"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000207,"raw_usage":{"total_tokens":1251,"prompt_tokens":769,"completion_tokens":482,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":405}},"tokens_in":513,"tokens_out":482,"duration_ms":4541,"temperature":1.0,"reasoning_tokens":405,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T01:11:36.771223+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the full 23-probe suite multiple times (for example, ten repetitions) on SHarD and on the best securely configured commercial agent under identical conditions. If SHarD's adjusted score falls below the commercial agent's, or if any category regresses relative to the unhardened harness, the equivalency claim fails. Alternatively, repeat the specific content-protection test that passed without the relevant control twenty times and count how often the agent spontaneously blocks the injection.","supporting_citations":[],"review_version":1}