{"id":"3868ac95-1243-4c5b-9d81-7d93e82302c6","arxiv_id":"2607.23438","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Agentic AI governance should assign Allowed Autonomy Levels separately from Autonomous Capability Levels so high capability need not imply high permission.","lead":"The paper separates what AI agents can do (capability levels) from what they are allowed to do (permission levels) for enterprise governance. Organizations can use the split to deploy capable agents under tighter human oversight when risk or reversibility demands it.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The only two demonstrations of the ACL axis are self-assessments by the framework's own authors, and they turn on exactly the subtle C3/C4 boundary call that the framework's consistency claim depends on; without independent-rater evidence, ACL may not be the stable, policy-independent upper bound th","rationale":"The reader's weakest_assumption — that ACL can be assessed consistently enough to serve as a stable upper bound, without multi-rater reliability or external validation — is the right target, and my read sharpens it: the problem is not merely the absence of reliability data but that the paper's only two data points are author self-assessments whose C3/C4 distinction rests on an architectural nuance (goal-chaining vs. intent-routing) that A.1 and B.2 describe in nearly identical language. This is the least secure load-bearing condition for the central claim, because the governance value of the AAL/ACL separation depends on ACL being a property others can reproduce, not one the designers can rationalize post hoc. I keep the verdict at CONDITIONAL rather than moving to REJECT for three reasons: (1) the conceptual separation itself is sound, clearly argued, and consistent with how human organizations decouple capability from authority — the paper's core contribution does not require the empirical premise to be already validated; (2) the paper is transparent about the limitation, listing calibration and cross-team consistency as future work, and the case studies are offered as illustrations, not as a benchmark; (3) the flaw is remediable with a well-defined study (the concrete test above), after which the framework would stand or fall on evidence rather than assertion. I also note a minor internal tension — Table 2's \"Highest AAL\" cap vs. the text's allowance that less-capable systems may receive greater freedom — which merits a clarifying sentence but does not rise to the level of the primary concern. The reader's CONDITIONAL verdict with medium correctness risk and high confidence remains appropriate; my contribution is to make the validation gap maximally concrete and falsifiable.","tokens_in":9893,"tokens_out":3158,"duration_ms":29823,"concrete_test":"Run a blinded multi-rater study: give 8–12 reviewers unaffiliated with the authors the Appendix A/B system descriptions (scrubbed of the paper's own C4/C3 verdicts) plus 3–5 additional agentic systems of varying design, and have each assign an ACL using only Table 2 and the level definitions. Report Fleiss' kappa and the agreement rate with the paper's C4/C3 labels. If kappa < 0.6, or if a majority of raters place the two case systems at the same level (erasing the chaining-vs-routing distinction), the \"stable upper bound\" premise fails as stated and ACL needs an operationalized checklist rubric before it can anchor AAL decisions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim rests on ACL being an \"inherent technical\" property, \"independent of organizational policy,\" assessed from observable behaviors, which then serves as a constraint on feasible AAL (decision framework, step 1: \"This assessment constrains the set of Allowed Autonomy Levels that are feasible\"). For that to function as governance infrastructure, different reviewers must land on the same ACL for the same system — the paper itself names this under its \"rigor\" principle (\"allow different reviewers, teams, and organizations to arrive at consistent conclusions\"). Yet the only empirical content in the paper is two systems assessed by the authors themselves, and the distinguishing judgment between them is delicate: Appendix A.2 rates the data-engineering system C4 because the supervisor chains gap-detection → diagnosis → remediation toward a compound goal, while Appendix B.2 rates the copilot C3 because its supervisor \"follows an established workflow and performs intent-based routing.\" But A.1 describes the C4 system's supervisor in the same terms — \"the supervisor classifies user intent and delegates tasks to four specialized sub-agents.\" The C3/C4 call therefore hinges on whether routing output feeds a subsequent agent, a fine architectural distinction that independent reviewers could easily score differently, especially since the paper concedes C1/C2 boundaries are \"blurred\" and \"on a continuous spectrum.\" A secondary, smaller tension: Table 2's \"Highest AAL\" column (C1→A1 … C5→A5) encodes a strict cap, while the decision-framework text allows that \"a less capable system may be granted greater freedom within a tightly constrained scope\" — the paper never reconciles whether ACL is a hard ceiling or advisory, and the cap direction is precisely what gives the separation its safety function. Nothing here is dishonest; the cases are plausibly labeled. But the load-bearing premise — ACL as a reproducible, reviewer-independent anchor — is asserted via self-asses","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper proposes a governance framework for agentic AI that separates two axes: Allowed Autonomy Levels (AAL A1–A5), the autonomy a system is authorized to exercise given risk, oversight, and accountability; and Autonomous Capability Levels (ACL C1–C5), its inherent technical abilities assessed from observable perception/reasoning/planning/tool-use/adaptation behaviors. A two-step decision process (Figure 1) first assesses ACL as an upper bound on feasible AAL, then selects AAL via an existing enterprise risk-assessment process, defaulting to A1. Tables 1–2 define the levels with human-role descriptors, accountability implications, and examples. Two deployed enterprise systems are presented as illustrations: a data-engineering multi-agent system rated C4 but deployed at A3, and an energy-operations copilot rated C3 but deployed at A2. The framework is positioned as an extension of the authors' NIST-aligned enterprise Responsible AI program.","tokens_in":10218,"tokens_out":2319,"duration_ms":106060,"significance":"If the framework holds up, the capability/permission separation is a genuinely useful conceptual contribution to a literature that mostly collapses the two (e.g., SAE-style ladders, human-in-the-loop taxonomies). The paper's strengths are practical rather than formal: it ships two detailed, real deployed-system assessments (Appendices A–B) that actually exercise the claimed separation in both directions of the delta (C4/A3 and C3/A2), the AAL progression is tied cleanly to an accountability narrative (errors-bounded → advice → actions → design errors → governance failures), and the decision process terminates at a well-defined floor (A1) with an explicit no-deploy condition. There are no fitted parameters or circular derivations; the taxonomy is stipulated and applied interpretively, which is normal for framework papers. The main limitation on significance is evidentiary: the entire empirical content is two self-assessments by the framework's own designers.","major_comments":[{"comment":"§Design Principles ('Rigor') vs. Appendices A.2/B.2: the framework's consistency claim — that ACL criteria 'allow different reviewers, teams, and organizations to arrive at consistent conclusions' — is load-bearing, because ACL functions in Figure 1 (step 1) as the constraint on feasible AAL, i.e., as governance infrastructure. Yet the only two ACL assessments in the paper are by the authors themselves, and they turn on exactly the delicate C3/C4 boundary. A.1 describes the C4-rated system's supervisor as one that 'classifies user intent and delegates tasks to four specialized sub-agents,' while B.2 rates the copilot C3 in part because its supervisor 'follows an established workflow and performs intent-based routing.' The operative distinction (chaining gap-detection → diagnosis → remediation toward a compound goal vs. dispatching each query to a single specialist) is a fine architectura","section":"Design Principles / Figure 1 / Appendices A.2, B.2"},{"comment":"The claim that the cases 'demonstrate the framework's effectiveness in managing risk' (§Example Applications) overstates what two self-assessed deployments can show. What the cases actually demonstrate is interpretability — that the AAL/ACL vocabulary can be applied coherently to real systems and that a capability–authorization gap is a deliberate, articulable governance choice. Nothing in the paper shows that the separation changed an outcome relative to the counterfactual (e.g., that without the framework these systems would have been deployed at A4/A3), nor that the AAL selections were correct. Please temper the effectiveness language or provide the counterfactual/governance-decision evidence (e.g., intake records showing the AAL assignment altered deployment controls).","section":"Example Applications (final paragraph)"},{"comment":"Table 2, 'Highest AAL' column: as written, this column hard-codes a 1:1 mapping (C1→A1, ..., C5→A5), which is in tension with the paper's central thesis that capability 'constrains... but does not determine' authorization and that 'a less capable system may be granted greater freedom within a tightly constrained scope.' The column is presumably meant as 'AAL may not exceed the AAL corresponding to ACL,' but it reads as a lookup table that re-couples the axes. Please clarify whether the mapping is a strict feasibility ceiling, and if so, state the rule explicitly (AAL ≤ f(ACL)) and reconcile it with the 'less capable system granted greater freedom' sentence in §Autonomy Level Decision Framework.","section":"Table 2"}],"minor_comments":[{"comment":"Reference verification: Wang et al. 2026, 'OpenClaw-RL: Train Any Agent Simply by Talking,' arXiv:2603.1016, could not be verified as a real arXiv identifier (the ID format is also malformed), and the AI Index 2026 citation is forward-dated. Please check all references for accuracy.","section":"References"},{"comment":"§Allowed Autonomy Level, opening paragraph: 'AAL concerns the level of accountability that an organization is willing to assume' conflates the definiens — AAL is elsewhere defined as authorized autonomy of action, with accountability as a consequence. Consider rephrasing so the level is defined by authorized action scope, with the accountability implication kept in Table 1's final column.","section":"Allowed Autonomy Level"},{"comment":"A3 definition bundles three disjunctive conditions ('require human approval, must be reversible... or must be accompanied by mitigation mechanisms'). The or-structure makes A3 assignment under-determined; a sentence on which condition is primary would help reviewers apply it.","section":"Allowed Autonomy Level, A3"},{"comment":"Figures A1/A2 are referenced as architecture diagrams but are central to the C3/C4 boundary argument; please ensure the routing-vs-planning distinction is visible in the figures themselves, not only in the prose.","section":"Appendix, Figures A1–A2"},{"comment":"Cohen et al. 2025 (Science) is cited in the reference list but does not appear to be cited in the body; either cite it in the related-work discussion or remove it.","section":"References"},{"comment":"Typo: 'risk assessment questionaries' → 'questionnaires' (§Autonomy Level Decision Framework).","section":"Autonomy Level Decision Framework"}],"recommendation":"major_revision","confidential_remarks":"The authors are from ExxonMobil and the two case studies are internal systems; the self-assessment limitation is therefore partly structural (external raters may not have access), but the paper should acknowledge this explicitly rather than presenting the rigor principle as satisfied. The venue fit is reasonable for a framework/practice paper. One reference (OpenClaw-RL) appears possibly fabricated or AI-generated; worth asking the authors to double-check the bibliography."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing is the clean separation: ACL (what the system can do) as an upper bound, AAL (what you authorize) chosen by risk, reversibility, and accountability. That decoupling is the contribution. Prior ladders (SAE, Sheridan–Verplank, Feng et al. human-involvement levels, liability taxonomies) mostly collapse capability and permission into one scale. This paper names both axes, gives a two-step flowchart, and shows two production multi-agent systems deliberately run below capability (C4→A3 data engineering; C3→A2 energy copilot).\n\nWhat it does well: the tables are usable, the informal labels (hands on / eyes on / mind on) travel in mixed rooms, and the cases are not toy demos. They actually exercise the claimed move—high technical reach, constrained authorization—for reasons that match the framework (irreversibility, downstream corruption risk, need for domain judgment). Integration with an existing NIST/EU-aligned RAI intake is practical rather than rhetorical. Circularity is low; they stipulate taxonomies and apply them.\n\nSoft spots, in proportion. The load-bearing claim is that ACL is a stable, policy-independent technical property different reviewers can score consistently. The only evidence is author self-assessment, and the C3/C4 call turns on a fine architectural distinction (intent routing vs chaining toward a compound goal) that independent raters could easily disagree on—especially since the paper already admits blurred boundaries lower down. Table 2’s “Highest AAL” column reads like a hard ceiling while the prose allows a less capable system more freedom in a tight scope; that tension is minor but unresolved. No multi-rater study, no external site, no outcome metrics on whether the gates change incidents or approval time. Free parameters (five levels, risk thresholds) are organization-specific by design.\n\nThis is for people building or reviewing enterprise agent governance, not for theory or benchmarks. Serious enough for peer review: clear thinking, honest about scope, real deployments. I would engage, reuse the AAL/ACL language, and push for reliability checks if it lands in a venue. Not a foundational result; solid applied framework work.","headline":"Useful dual-axis split (capability vs permission) with two real deployments; the idea is clear and operational, the validation is thin and self-scored.","tokens_in":11569,"tokens_out":550,"would_cite":true,"duration_ms":19722,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Agentic AI governance works only when what a system can do is separated from what it is allowed to do.","keywords":["agentic AI","autonomy levels","Allowed Autonomy Levels","Autonomous Capability Levels","AI governance","risk-aware deployment","human oversight","enterprise AI"],"falsifier":"Have independent governance teams rate the same multi-agent systems on ACL and AAL; if ratings diverge widely, or if systems deployed under the framework still escalate beyond approved autonomy or fail risk controls no better than mixed single-scale taxonomies, the separation does not operationalize as claimed.","tokens_in":11220,"feed_emoji":"⚖️","tokens_out":835,"duration_ms":18843,"temperature":0.7,"pith_summary":"This paper argues that talks about AI autonomy usually mix two different things: technical ability and permission to act. It offers a practical governance split—Autonomous Capability Levels (what an agent can do) versus Allowed Autonomy Levels (what an organization authorizes it to do given risk, oversight, and accountability). Five paired ladders run from reactive tools and hands-on human control up to independent operations and delegated authority, and a risk-aware decision process picks the allowed level only after capability is assessed. Two deployed enterprise multi-agent systems show the point in practice: high-capability agents were deliberately held one step below their technical ceiling because of reversibility, safety, and organizational readiness. The claim is that this separation lets organizations move faster and safer without treating every capable agent as automatically free to act.","feed_headline":"Separate what AI agents can do from what they may do","feed_subtitle":"A five-level capability and permission split lets enterprises constrain high-skill agents by risk","key_machinery":"The AAL/ACL separation plus a two-step decision process: first assess ACL from observable behaviors, then assign AAL only if risk is acceptable under existing responsible-AI controls, with A1 as the floor and redesign required if even A1 is unacceptable.","core_discovery":"The central claim is that Allowed Autonomy Levels (AAL)—authorization under risk, oversight, and accountability—must be decided separately from Autonomous Capability Levels (ACL)—inherent technical abilities in perception, reasoning, planning, tool use, and adaptation. Capability sets an upper bound; risk, reversibility, and organizational readiness choose the permitted level, so a high-ACL system can and often should run at a lower AAL.","pith_inferences":["Regulators and insurers could map liability and audit requirements to AAL rather than to model size or benchmark scores, treating capability as evidence of possible harm only.","Product teams may need explicit ‘autonomy dials’ in architecture—approval gates, reversible actions, kill switches—so the same ACL stack can be shipped at multiple AAL settings.","If ACL scoring proves inconsistent across raters, the framework’s next empirical test is inter-rater reliability and calibration guides, not more level labels."],"forward_implications":["Enterprise reviews can approve capable agents at deliberately lower autonomy without re-scoring technical ability when policy tightens or loosens.","Autonomy can be raised or lowered incrementally as reversibility, audit trails, and organizational readiness improve, without reinterpreting capability.","Most production agentic apps are expected to sit at A1–A3; A4–A5 require stronger upfront justification and safeguards.","Multi-agent systems must be capability-rated as composed systems, not as isolated sub-agents, because autonomy emerges from orchestration."],"fun_headline_variants":["Separate AI capability from allowed autonomy","Authorize agent autonomy by risk not skill","High-capability agents can run at lower AAL","Capability sets the ceiling permission sets the level","Govern what agents may do apart from what they can"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That a composed multi-agent system’s technical capability can be scored as a stable, policy-independent upper bound that different reviewers will assign consistently enough to govern real deployments.","fun_headline_variants_meta":{"raw":{"variants":["Separate AI capability from allowed autonomy","Authorize agent autonomy by risk not skill","High-capability agents can run at lower AAL","Capability sets the ceiling permission sets the level","Govern what agents may do apart from what they can"]},"model":"grok-4.5","effort":"low","cost_usd":0.002618,"raw_usage":{"total_tokens":976,"prompt_tokens":755,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":26184000,"prompt_tokens_details":{"text_tokens":755,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":171,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":755,"tokens_out":50,"duration_ms":4486,"temperature":1.0,"reasoning_tokens":171,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T22:04:11.600384+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Have independent governance teams rate the same multi-agent systems on ACL and AAL; if ratings diverge widely, or if systems deployed under the framework still escalate beyond approved autonomy or fail risk controls no better than mixed single-scale taxonomies, the separation does not operationalize as claimed.","supporting_citations":[],"review_version":1}