{"id":"e404a757-2853-46f6-8c66-cd7337518ba3","arxiv_id":"2607.27288","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"OSB proposes frozen synthetic-enterprise snapshots with gold posture answers so AI agents can be benchmarked on security investigation via SQL or native vendor APIs.","lead":"Open Security Benchmark is a proposed framework for testing AI security agents on enterprise defense work: agents investigate frozen, fake company environments and are scored against known-correct answers. It matters because there is currently no public, standard way to check whether a security agent can be trusted on real defensive tasks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Primary metric relies on uncalibrated LLM judges despite 'closed-form ground truth' claim; without calibration data, OSB scores are judgments, not measurements.","rationale":"The reader correctly identified the transfer assumption as the weakest link for real-world relevance, and the paper's own §7 limitation statement supports that. However, the most load-bearing concern for the paper's central 'measurement rather than judgment' assertion is internal: the primary scoring pathway is an LLM judge, not a closed-form check. The paper's abstract and §1 emphasize closed-form ground truth, but §4.3 reveals that free-form answers are mapped onto that reference by a rubric-locked LLM panel, with expert calibration promised but not delivered. This means the claimed reproducibility and objectivity of OSB scorecards is not yet established, even for synthetic environments. This concern is concrete and testable, and it reinforces the CONDITIONAL verdict: the framework has a sound architectural design, but its central measurement guarantee requires shipping expert calibration data, inter-judge agreement statistics, and judge-stability analysis. Without those, the framework is a well-specced proposal, not a validated benchmark. I therefore agree with the reader's CONDITIONAL verdict, but for a reason that precedes and compounds the transfer concern: unreliable measurement makes transfer moot. The proposed calibration study is a single check that would settle whether the primary metric can support the paper's headline claim.","tokens_in":15025,"tokens_out":2812,"duration_ms":26869,"concrete_test":"Run a calibration study on the two identity packs: sample 100 agent outputs (varied tasks, scales), have 3 independent expert security analysts score each against the ground-truth atom, then run the proposed two-judge LLM panel with mode aggregation and low-grade tie-break. Compute Cohen's kappa between the panel and expert consensus, report a confusion matrix, and test the panel's sensitivity to benign paraphrasing (rephrasing correct answers, reordering subject lists, using synonyms for configuration facts). If kappa < 0.7, or >10% of expert-correct answers receive a sub-1.0 judge score, the primary metric is not calibrated and the 'measurement rather than judgment' claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"OSB's central claim (§1, §8) is that freezing an environment and grading against a 'closed-form ground truth' makes agent trustworthiness 'a matter of measurement rather than judgment.' But the primary metrics in Table 1 — AnswerCorrectnessVsGT, AnswerVerdict, ReasoningUtility, and SQLSemanticAppropriateness — are all scored by a rubric-locked LLM-as-judge panel, not by deterministic comparison. Only the structural Tables/Joins metrics are deterministic. The agent answers in free-form prose, so the judge must map that prose onto the ground-truth atom, a task that is inherently subjective and model-dependent. The paper acknowledges this by promising 'expert calibration' as the anchor (§4.3) and reporting inter-judge agreement, but it ships no calibration data, no kappa values, no confidence intervals. Without such evidence, scores can be systematically biased, drift across judge versions, or be fooled by paraphrases that preserve meaning but change wording. This is not an external validity issue — it undermines the framework even on the synthetic environments it ships, and it directly contradicts the 'measurement rather than judgment' conclusion. The transfer-to-real-tenants assumption (the reader's weakest assumption) is important, but it is secondary: if the primary metric is not a validated measurement, then even perfect transfer would still yield unreliable scores.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Open Security Benchmark (OSB), a framework for evaluating agentic AI on enterprise security-posture investigation. It defines two investigation modalities — text-to-SQL over a frozen relational snapshot and native-vendor API calls over a served instance of the same environment — and five framework components: data layer, task/evaluation layer, scoring layer, harness, and bring-your-own path. The framework is instantiated with two identity-security packs totaling 127 tasks and synthetic-organization datasets at three scales. The central claim is that OSB 'closes the environment data gap' by surfacing curated, frozen, synthetic enterprise environments with closed-form ground truth, making agent trustworthiness 'a matter of measurement rather than judgment.'","tokens_in":15286,"tokens_out":3277,"duration_ms":29768,"significance":"If fully realized, OSB would be a valuable shared substrate: the two-modality design is thoughtful, the minimal auditable harness with an evidence archive is a genuine strength, and the explicit trust boundary between agent-visible and evaluator-only artifacts is a sound methodological commitment. The paper also productively distinguishes deterministic structural metrics from semantic judged metrics, and the plan to report bootstrap confidence intervals and inter-judge agreement is appropriate. However, the central claims are currently conditional on artifacts that are not shipped, and the primary semantic metrics are rubric-locked LLM judgments with no calibration evidence. The strengths are real but prospective; the paper is a detailed framework proposal rather than a demonstrated benchmark.","major_comments":[{"comment":"The load-bearing claim that OSB 'closes the environment data gap' is unsupported as stated. Section 5 says the datasets 'will be released' and 'the initial catalog is in preparation,' and Table 3 describes 'the reference snapshot' rather than a published artifact. No dataset identifier, hash, or download is provided. Contribution (ii), the curated data catalog, is therefore unverifiable; the central evaluation cannot be reproduced or independently checked. Shipping the datasets and code is a prerequisite for the paper's main claim.","section":"§5, §1"},{"comment":"The primary metrics are not closed-form measurements. AnswerCorrectnessVsGT, AnswerVerdict, ReasoningUtility, and SQLSemanticAppropriateness are all judged by an LLM panel that maps free-form agent prose onto the 'ground-truth atom'; only Tables/Joins recall/precision/F1 are deterministic. The paper states that expert calibration 'is to be validated' and reports no calibration data, no Cohen's kappa, no confidence intervals, and no judge-bias analysis. This directly contradicts the §8 statement that trust becomes 'a matter of measurement rather than judgment.' The scoring method needs a calibration study on expert-labeled outputs before the primary metrics can be treated as measurements.","section":"§4.3, Table 1"},{"comment":"The evaluation chain is self-referential by the paper's own account. Section 4.3 says the scoring criteria are 'adopted from the cross-vendor ISPM benchmark [30], reusing its Appendix B rubric texts verbatim,' and Section 6 says both packs are 'derived from the two public Sola ISPM benchmarks [29, 30].' Since these are the same authors' prior benchmarks, the framework's validity depends entirely on the quality of those unpublished peer benchmarks. The paper needs to provide independent evidence that the tasks and rubrics are valid for posture investigation, or at least a clear comparison showing what OSB adds beyond repackaging prior work.","section":"§4.3, §6"},{"comment":"The synthetic-to-real transfer assumption is acknowledged but not supported. Section 5 asserts the environments 'are synthesized from real ones' with production data patterns, but no data distributions, realism metrics, or comparison to production telemetry are given. Section 7 concedes that fabricated organizations 'do not reproduce a live tenant's operational drift and incident telemetry.' Without a validation protocol or any evidence of transfer, the claim that OSB scores predict performance on real deployments remains an unstated assumption rather than a demonstrated property.","section":"§5, §7"},{"comment":"The native-API modality is claimed as a full second surface, but the manuscript says only 'a subset of the eight sources is served this way today' without stating which subset, and the structural role of the request trace is described qualitatively rather than defined as a concrete metric. If the native-API modality is part of the central contribution, it needs a definite specification and a working implementation for all claimed sources.","section":"§5.2"}],"minor_comments":[{"comment":"The table headers contain apparent typos ('F amily', 'V endor') and should be cleaned before publication.","section":"Table 1"},{"comment":"The Okta per-application 'mfa_required' simplification is disclosed, which is good, but the paper should explicitly discuss whether this simplification makes the corresponding tasks unrepresentative of real Okta configuration models and whether the benchmark's realism claim depends on such simplifications.","section":"§6"},{"comment":"The description of the judge panel is underspecified: 'several independent traces' is not quantified, and the tie-breaking rule 'toward the lower grade' could systematically bias scores downward. A precise protocol is needed for reproducibility.","section":"§4.3"},{"comment":"The three visibility classes (open, gated, private) are described, but only open/private ship today; the gated leaderboard is future work. The paper should label the current status more explicitly in the framework description so readers do not infer a fully implemented gated mechanism.","section":"§4.2"},{"comment":"The evidence archive caveats are appropriately honest, but 'a stored run can be reviewed again but not yet automatically re-scored' is a significant limitation for the claimed training-loop utility; this should be stated earlier and more prominently.","section":"§7"}],"recommendation":"major_revision","confidential_remarks":"This is a framework proposal from the vendor's own research team, and its two packs and scoring rubrics are drawn verbatim from the same team's prior benchmarks. The central empirical claims are contingent on releasing the datasets and on a calibration study that is not yet performed. I would treat this as a major revision rather than a reject: the framework design is coherent and the limitations are honestly disclosed, but the paper as submitted does not yet substantiate its headline claims. The editor may want to require artifact submission and an independent calibration study as conditions of acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a framework paper, not a benchmark release: the datasets are in preparation and there are no baseline runs, commit hashes, or evaluation results. Second, the framework design is genuinely thoughtful and better specified than most: five components, two investigation modalities over the same frozen snapshot, plan-independent gold atoms, a trust boundary, and a BYO path for private tenants. That integration is new in the security-agent evaluation space, and the authors position it carefully against Spider, BIRD, Cybench, OrgAccess, and configuration-as-SQL tools. The soft spots are real but mostly conditional. The closing-the-environment-data-gap claim in Section 1 and Section 8 is unsupported by anything shipped. That needs to be fixed before this is a benchmark rather than a proposal. More structural: the primary metrics in Table 1 (AnswerCorrectnessVsGT, AnswerVerdict, ReasoningUtility, SQLSemanticAppropriateness) are all LLM-judge rubrics. The paper calls the ground truth closed-form and says trust becomes a matter of measurement rather than judgment, but the scores are judgments until the promised expert calibration data actually appears. The deterministic Tables/Joins metrics are genuinely measurement, but they are secondary. This is not just an external-validity issue; it affects the synthetic environments too. Third, the validation chain is self-referential: tasks are derived from the authors own Sola ISPM benchmarks and the rubric is reused verbatim from reference 30. That does not make it wrong, but it does mean content-level validity currently reduces to the authors prior work. Finally, the synthetic-to-real transfer assumption is acknowledged in Section 7, and it is the right caveat; it just has not been tested. None of this kills the proposal. The architecture is clear enough that a referee could meaningfully push on it, and the authors are unusually honest about what is missing. Who should read it: anyone building or evaluating defensive security agents, and anyone working on LLM-as-judge methodology. It deserves a serious referee, conditionally, with the expectation that artifacts and calibration data arrive in revision. I would take it to a reading group now; I would cite it once the data is out.","headline":"Well-specified framework for a real gap, but the 'closes the gap' claim outruns what ships; needs artifacts and judge calibration before it is a benchmark rather than a proposal.","tokens_in":724,"tokens_out":836,"would_cite":true,"duration_ms":28576,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A frozen, synthetic enterprise environment with gold answers can turn AI posture-agent trust into a measurable quantity.","keywords":["autonomous cyber defense","security posture management","agentic AI","benchmarking","text-to-SQL","identity security","synthetic environments","LLM evaluation"],"falsifier":"Have human analysts score the same agents on a live (or realistically drifted) enterprise environment and compare the rank order with the agents' OSB scores; if the rankings diverge substantially, the synthetic-to-real transfer fails.","tokens_in":14796,"feed_emoji":"🛡️","tokens_out":6778,"duration_ms":50739,"temperature":0.7,"pith_summary":"The paper argues that the field cannot yet answer whether an AI security agent should be trusted, because real enterprise environments are private, cross-vendor, and correlated, and no shared, queryable target exists for evaluating posture-investigation agents end to end. To close this 'environment data gap,' the authors propose Open Security Benchmark (OSB), which freezes synthetic-but-realistically-shaped enterprise environments into immutable relational snapshots with closed-form gold answers, and evaluates agents either by text-to-SQL over the snapshot or by native API calls against a served instance of the same data. Because the ground truth is fixed and scoring is answer-based, different query strategies for the same question score alike, and because both modalities read from the same snapshot, they measure the same work. If OSB works as claimed, practitioners and researchers can compare posture agents on identical, rotating environments, and trusting a posture agent becomes a matter of measurement rather than judgment.","feed_headline":"Frozen synthetic enterprises let you measure AI security agents","feed_subtitle":"A shared, immutable snapshot with gold answers replaces guesswork about which posture agent to trust.","key_machinery":"The load-bearing object is the frozen environment snapshot: a read-only relational database of a synthetic enterprise—44 tables spanning eight vendor-style sources at three organizational scales—bundled per task with a 'ground-truth atom' that fixes the verdict, the closed set of matching subjects, the determining configuration facts, and the minimal sufficient tables/joins. This snapshot is what makes scoring plan-independent, because any SQL or API trace that produces the same answer atom is scored alike. Around it, the framework's harness exposes only two tools (get-schema and run-query) and records every action into an evidence archive, while the scoring layer runs deterministic structur","core_discovery":"The central claim is that a frozen, content-addressed enterprise environment—a synthetic organization whose vendor tables and data patterns mirror real product data models—can serve as a closed-form ground truth for scoring agentic AI on security posture investigation. The authors assert that this environment can be presented through two interrogation surfaces: a relational text-to-SQL snapshot and a natively served API instance, both backed by the same rows so that a finding is defined once and posed either as a query workload or as a sequence of API calls. Around this frozen snapshot sits a five-component framework: a data layer, a task/evaluation-set layer with a ground-truth atom (verdic","pith_inferences":["Editorial inference: If the synthetic-to-real transfer holds, OSB could serve as a standardized pre-deployment check for security agents in regulated industries that cannot share live tenant data, giving buyers a defensible basis for procurement decisions.","Editorial inference: The native-API modality implies that a vendor could benchmark agents against its own API surface without exposing a consolidated schema, potentially turning the framework into a venue for vendor-specific agent evaluations.","Editorial inference: Because scoring is answer-based, any future interrogation surface—graph queries for transitive access paths, CLI command sequences—would plug into the same ground truth, so the framework's coverage could grow without re-anchoring answers.","Editorial inference: A testable extension would be to publish the promised expert-calibration agreement statistics; that would let researchers quantify how much judge-panel drift limits score comparability across model configurations."],"forward_implications":["Posture-investigation agents can be compared on identical, immutable enterprise environments, so score differences are attributable to the agent rather than the data.","The same snapshot and answer criteria make text-to-SQL and native-API results directly comparable, so a model that passes in one modality can be checked in the other.","Rotating the environment (regenerating and republishing a new content-addressed revision) keeps answer-level memorization from carrying over to the next revision.","The multi-axis scorecard (answer correctness, verdict, reasoning utility, SQL quality, structural table/join recall) lets teams localize whether an agent fails on comprehension, correlation, or planning.","Every run's evidence archive is contract-conformant training data, so the benchmark's output can feed a training loop that improves the next agent rather than only scoring it."],"fun_headline_variants":["Frozen enterprise snapshots give AI security agents a scored test","New benchmark scores AI security agents on frozen enterprise data","OSB freezes enterprise state to give agentic AI an auditable test","Frozen synthetic orgs give ground truth for AI security agents","Measure trust in AI security agents with frozen enterprise snapshots"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That scores on synthetic enterprises synthesized from real product data models predict an agent's performance on a live tenant's environment with its operational drift and incident telemetry—a transfer the paper does not yet demonstrate.","fun_headline_variants_meta":{"raw":{"variants":["Frozen enterprise snapshots give AI security agents a scored test","New benchmark scores AI security agents on frozen enterprise data","OSB freezes enterprise state to give agentic AI an auditable test","Frozen synthetic orgs give ground truth for AI security agents","Measure trust in AI security agents with frozen enterprise snapshots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000793,"raw_usage":{"total_tokens":3367,"prompt_tokens":818,"completion_tokens":2549,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":2463}},"tokens_in":562,"tokens_out":2549,"duration_ms":15055,"temperature":1.0,"reasoning_tokens":2463,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T10:20:03.916878+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human analysts score the same agents on a live (or realistically drifted) enterprise environment and compare the rank order with the agents' OSB scores; if the rankings diverge substantially, the synthetic-to-real transfer fails.","supporting_citations":[],"review_version":1}