{"paper":{"title":"Cochise: A Reference Harness for Autonomous Penetration Testing","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Cochise supplies a minimal 597-line Python harness that separates planning from execution to support reproducible comparisons of LLM agents in penetration testing.","cross_cats":["cs.AI","cs.SE"],"primary_cat":"cs.CR","authors_text":"Andreas Happe, J\\\"urgen Cito","submitted_at":"2026-05-12T07:28:12Z","abstract_excerpt":"Recent work on LLM-driven autonomous penetration testing reports promising results, but existing systems often bundle architectural, prompting, and tool-integration choices together. This makes it difficult to determine what is gained over a simple agent and harness. We present Cochise, a 630 LOC Python reference implementation for autonomous penetration-testing experiments. Cochise connects to a Linux execution host over SSH and supports attacking controlled target environments reachable from that jump host.\n  The prototype implements a Planner--Executor architecture in which long-term state "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Cochise is intended not as a state-of-the-art pen-testing agent, but as reusable experimental infrastructure for comparing models, agent architectures, and penetration-testing traces.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That a minimal separated Planner-Executor architecture connected via SSH to a Linux host is sufficient to enable meaningful comparisons of LLM choices and agent designs across different penetration-testing scenarios.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Cochise is a minimal open-source harness implementing a Planner-Executor architecture over SSH for LLM-driven autonomous penetration testing, demonstrated on the GOAD testbed with released replay tools and trajectory logs.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Cochise supplies a minimal 597-line Python harness that separates planning from execution to support reproducible comparisons of LLM agents in penetration testing.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"3e47aea00163169b0a38934adf42f41ef89cf8634cc8298827cbdf1e58e6230f"},"source":{"id":"2605.11671","kind":"arxiv","version":2},"verdict":{"id":"4d33dc8e-668b-4986-abe8-e0d2451c9acf","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-13T01:02:50.261474Z","strongest_claim":"Cochise is intended not as a state-of-the-art pen-testing agent, but as reusable experimental infrastructure for comparing models, agent architectures, and penetration-testing traces.","one_line_summary":"Cochise is a minimal open-source harness implementing a Planner-Executor architecture over SSH for LLM-driven autonomous penetration testing, demonstrated on the GOAD testbed with released replay tools and trajectory logs.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That a minimal separated Planner-Executor architecture connected via SSH to a Linux host is sufficient to enable meaningful comparisons of LLM choices and agent designs across different penetration-testing scenarios.","pith_extraction_headline":"Cochise supplies a minimal 597-line Python harness that separates planning from execution to support reproducible comparisons of LLM agents in penetration testing."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2605.11671/integrity.json","findings":[],"available":true,"detectors_run":[{"name":"doi_title_agreement","ran_at":"2026-05-21T00:01:32.354152Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_compliance","ran_at":"2026-05-20T14:07:12.726511Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"claim_evidence","ran_at":"2026-05-20T03:42:00.541479Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"ai_meta_artifact","ran_at":"2026-05-19T11:40:09.424891Z","status":"completed","version":"1.0.0","findings_count":0}],"snapshot_sha256":"2f56123c23f7f752107ae214c0e25b40aa9c93e5ca440ea8803342678c720774"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}