{"id":"beb01d46-f9f5-4a79-b07a-48db8a02f062","arxiv_id":"2607.18970","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Agent skills should be managed as persistent software units with identity, lifecycle, and engineering structure, not just prompt files.","lead":"This paper defines 'Skillware': agent skills treated as real software objects with identity, versioning, and a lifecycle. It gives the AI-agent world a shared language for building, maintaining, and evolving reusable behaviors.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"C1 'behavioral primacy' is a subjective judgment with no inter-rater reliability; if independent coders cannot reproduce the C1–C3 classifications, the claim that Skillware is an auditable, recurring software surface is not empirically established.","rationale":"The paper's central claim is a coherent ontology, and the authors are admirably explicit about their evidence boundaries: the corpus is registry-heavy, the cases are purposive, no inter-rater reliability is reported, and several pattern mappings are constructive rather than observed. The frozen revisions, content hashes, manifest release, and independent studies are real supporting evidence. The remaining load-bearing concern is that C1—the condition that separates Skillware from adjacent artifacts—is an interpretive judgment about 'primacy' that static inspection cannot settle. Because all positive classifications were made by the authors, the possibility of systematic idiosyncrasy is not just a sampling issue; it threatens the operational auditability that the definition promises. However, the reader already assigned CONDITIONAL on essentially these grounds, and my concern does not show the claim is false or internally inconsistent. It shows that the empirical hinge needs independent validation before the universality of the ontology can be accepted. Therefore the appropriate verdict remains CONDITIONAL, and no change to the reader's verdict is needed.","tokens_in":17540,"tokens_out":4102,"duration_ms":45205,"concrete_test":"Recruit two coders with access only to the §3.3 protocol and the public evidence supplement. Have them independently classify, at frozen revisions: (a) all 15 boundary cases and 13 technical cases; and (b) 50 randomly sampled repository IDs from the SkillMD-138K manifest with their resolved SKILL.md directories. Each coder assigns C1, C2, C3, and Lifecycle Continuity. Compute Cohen's kappa (or Krippendorff's alpha) for C1 and for overall C1–C3 membership. If alpha < 0.6 for C1 or for membership, the category boundary is not reliably operational and the empirical support for a recurring Skillware surface is weakened. If alpha ≥ 0.8, the concern is answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 makes category membership auditable through C1–C3, and the central conclusion in §7.4 rests on these conditions picking out a stable, existing category. The weakest load-bearing premise is C1, 'Skill-centered behavioral primacy': it asks an analyst to decide which artifact is 'primary' in a unit's behavior, routing, and user-facing identity. The paper itself concedes in §7.2 that primacy can be hard to judge in packages with substantial executable services and that different analysts may select different unit boundaries. Yet all 15 boundary cases and 13 technical cases in §4.4 are purposively selected by the authors, and §4.4 explicitly states: 'It supplies no estimate of inter-rater reliability.' The constructive pattern fixtures in §5.3 are authored by the same team, so they show constructibility, not independent recurrence. If C1 is unavoidably a subjective judgment about authorial intent that static source inspection cannot verify, then the 12 inclusions and 3 negatives may reflect the authors' prior commitment to the ontology rather than an observed category boundary. That would not make the conceptual proposal false, but it would sever the empirical support for 'recurring artifact envelope' and 'auditable membership' from the central claim. This is the same core weakness the reader identified, and the paper's own limitation statements do not repair it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Skillware, a software ontology and engineering lifecycle for persistent behavioral artifacts (Agent Skills) in AI agent systems. It defines a Behavioral Artifact and a Skillware Unit, with three necessary conditions for category membership (C1 behavioral primacy, C2 independent software identity, C3 Agent Host execution relationship), plus a separate Lifecycle Continuity property. Evidence combines the Agent Skills specification, a frozen corpus of 138,133 SKILL.md records, independent empirical studies, 15 purposively selected category-boundary cases, 13 fixed-revision technical cases, and constructive design-pattern fixtures. The central claim is that persistent natural-language behavioral specifications create a managed software surface, and Skillware is the abstraction that provides addressable artifacts, independent identity, activation relations, engineering structure, and a basis for identity-preserving change.","tokens_in":17898,"tokens_out":3081,"duration_ms":25823,"significance":"If the central claim holds, the paper supplies a useful definitional anchor for a fast-moving area: it makes a coherent proposal for what counts as a Skillware Unit, separates structural from lifecycle concerns, and connects Agent Skills to established software-engineering concepts. The paper is generally careful to bound its claims (no prevalence claims, no behavioral-equivalence claims, explicit inference limits per evidence layer). Concrete strengths are the frozen corpus with pinned hashes and repository revisions, the auditable case-review protocol with counterevidence fields, and the explicit admission of limitations in §7.2. The definitional contribution is plausible and could support future empirical and theoretical work. The main risk is that the empirical support for the 'recurring artifact envelope' is weaker than the presentation sometimes implies, because the case classifications rely on a subjective C1 judgment with no inter-rater reliability.","major_comments":[{"comment":"The load-bearing empirical claim that C1–C3 identify a stable, auditable category rests on purposively selected cases. §4.4 explicitly states 'It supplies no estimate of inter-rater reliability,' and §7.2 concedes that primacy can be difficult to judge in packages with substantial executable services and that different analysts may select different unit boundaries. Since C1 is a judgment about which artifact is primary in behavior, routing, and user-facing identity, static source inspection alone does not verify it. The paper's own limitation statements do not repair this. I do not think this invalidates the conceptual proposal, but it does mean the 'empirical evidence establishes a recurring artifact envelope' framing in §7.4 overstates what the case review supports. The authors should either (a) rescope the claim to 'the case review demonstrates that the category can be operationalized","section":"§4.4, §7.2"},{"comment":"The design-pattern transfer evidence is presented as software-continuity evidence, but the main-text mappings combine the authors' own constructive fixtures (e.g., the release-subject fixture in §5.3.4, the evidence-strategy fixture in §5.3.6) with selected open-source correspondences. Table 5 itself marks most mappings as 'candidate' or 'constructive.' This shows constructibility, which is a legitimate contribution, but it does not establish recurrence in the ecosystem. The paper mostly says this ('Ecosystem frequency and comparative benefit remain empirical questions'), so the main issue is presentation: the conclusion (§7.4) lists 'documented or reconstructed activation relations' as evidence for the central claim. Recommended fix: separate 'constructibility' from 'recurrence' more sharply in the abstract and conclusion.","section":"§5.3, Table 5"},{"comment":"The 138,133-record corpus is registry-heavy and used for narrow claims, which the paper acknowledges. However, the abstract and §7.4 say the evidence 'establishes a recurring artifact envelope' and 'engineering pressure.' The corpus signals (98.73% frontmatter, 23.2% path tokens, median 169 lines) show that SKILL.md files have a structured envelope and explicit references, but they cannot establish the C1–C3 conjunction for any unit. Given that Definition 1's membership requires all three conditions, the leap from 'structured SKILL.md files exist at scale' to 'Skillware exists at scale' should be explicitly qualified in the conclusion. This is a framing rather than a correctness issue, but it matters because readers may attribute more empirical support to the central claim than the evidence design warrants.","section":"§4.2, Table 3"}],"minor_comments":[{"comment":"The definition has a mild circularity risk: it defines Skillware in terms of 'primary Behavioral Artifact' and C1 operationalizes primacy. The paper's operational tests help, but it would be cleaner to state upfront that the definition is an analytical stipulation with an empirical operationalization, not a discovered natural kind. This is already implicit in §7.2, but a one-sentence clarification at Definition 1 would help.","section":"§3.3, Definition 1"},{"comment":"The quoted Facade rule ('1% chance... MUST invoke') is used as evidence of a declarative operation. It could also be read as an over-triggering prompt-injection risk. A brief note acknowledging this as a design tradeoff would prevent the reader from mistaking the example for a best practice.","section":"§5.3.1"},{"comment":"The distinction between Self-Evolution (candidate generation) and Collaborative Evolution (governance topology) is clear, but §6.3's 'Deployment and Field Adaptation Hypothesis' is explicitly prospective. This is fine, but the section placement between evolution theory and limitations makes it easy to over-read as a finding. Suggest labeling it as 'hypothesis' in the section header.","section":"§6.2–6.3"},{"comment":"The evidence supplement is described as containing the boundary review and technical review, but the paper does not state how many coders performed the classifications or how disagreements would be resolved. Even reporting the coding procedure in detail would help auditability and would partially address the inter-rater concern.","section":"§4.1"},{"comment":"The paper uses 'skills-compatible agent' from the Agent Skills specification for C3. It would be helpful to clarify whether this is a formal compatibility certification or simply a working discovery/activation capability.","section":"General"}],"recommendation":"minor_revision","confidential_remarks":"The paper is a definitional/conceptual contribution with an empirical shell. The reader's conditional verdict and the skeptic's concern about C1 subjectivity are the key issues. I do not believe they warrant rejection because the paper explicitly discloses the limits and the ontology is coherent regardless of inter-rater reliability. However, the conclusion's language slightly overstates the empirical support. A minor revision that rescopes those sentences and adds a short paragraph on the coding procedure would be sufficient. I would also suggest the editors consider whether the design-pattern supplement is essential to the main claim; if not, the main text could be tightened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead Skillware. The useful part is the conceptual apparatus: distinguishing Skill Artifact from Skillware Unit, defining C1–C3 membership conditions, and separating Lifecycle Continuity as an orthogonal property. That gives researchers a shared vocabulary for versioning, maintenance, and governance of agent skills, and it connects to existing SE concepts without claiming word priority. The design-pattern mappings are also genuinely new — Facade in Superpowers, Adapter in gstack — and the seven-element admission protocol is a reasonable guard against name-based pattern claims. The frozen corpus (138K SKILL.md files) and pinned revisions are real evidence, and the paper is careful about what the corpus can't show: path tokens don't prove runtime use, frontmatter doesn't prove semantics.\n\nThe soft spot is exactly where the stress test lands: C1 behavioral primacy requires a judgment about which artifact is primary, and all positive and negative cases were selected by the authors. No inter-rater reliability. That means the \"recurring artifact envelope\" is demonstrated for the authors' reading, not for independent coders. The paper concedes this in §4.4 and §7.2, so it's an acknowledged limitation, but it still limits what the empirical sections can establish. The design-pattern fixtures in §5.3 are also authored by the same team; they show constructibility, not independent recurrence. I'd call that a moderate weakness, not a fatal one, because the central claim — that persistent NL behavioral specs form a software surface worth managing — doesn't depend on the empirical boundary cases. C1–C3 are stipulated; the surprise would be if the ontology couldn't classify real cases, and the case review suggests it can, at least for the authors.\n\nThe citation pattern looks fair: the paper positions against SkillOS, SkillFab, and the surveys, and it doesn't overclaim novelty. No free parameters, no invented entities beyond the named constructs.\n\nWho's it for: people working on agent-skill versioning, runtime governance, or skill ecosystems. A serious referee should engage — the conceptual work is strong enough to merit discussion, and the empirical gaps are fixable with independent coding of the boundary cases and a few more negative cases. I'd send it to review, and I'd cite it for the ontology even while noting the evidence limits.\n\nRecommendation: send to peer review, with the inter-rater reliability issue as the main requested revision.","headline":"The paper's real contribution is a workable ontology and lifecycle for treating agent skills as software; the empirical boundary evidence is thinner than the conceptual frame, but the authors say so themselves.","tokens_in":18309,"tokens_out":1510,"would_cite":true,"duration_ms":15359,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper defines Skillware, a software abstraction that lets persistent natural-language agent capabilities be treated as identifiable, composable, versioned software units with their own lifecycle.","keywords":["Skillware","Agent Skills","behavioral artifacts","natural-language programming","software ontology","software engineering lifecycle","identity-preserving evolution","design pattern transfer"],"falsifier":"Have independent raters apply C1-C3 to a random, revision-frozen sample of skill-bearing repositories, including ones with heavy executable support; if a substantial share fail C2 (no independently trackable identity) or C3 (no skills-compatible agent host activation path), or if raters disagree systematically on C1, then the claim that persistent behavioral artifacts form a recurring managed software surface would be refuted.","tokens_in":17452,"feed_emoji":"🧩","tokens_out":4557,"duration_ms":39679,"temperature":0.7,"pith_summary":"The paper tries to establish that agent skills—persistent natural-language task specifications that agents load and execute—have become a distinct kind of software object, and that software engineering can be extended to them through a named abstraction, Skillware. It defines three auditable conditions for membership—behavioral primacy, independent software identity, and an agent-host execution relationship—plus a separate property, Lifecycle Continuity, for whether the same unit survives updates, rollbacks, and removal. The authors argue that a recurring artifact envelope and engineering pressure already exist, evidenced by a frozen corpus of over 138,000 skill files, independent studies, and fixed-revision case reviews. If right, this gives agent capabilities an address, version, provenance, and change process, making them composable and maintainable like conventional software components.","feed_headline":"Three conditions turn agent skills into managed software","feed_subtitle":"Persistent natural-language behavior gains an address, identity, and versioned lifecycle for agent capabilities.","key_machinery":"The central machinery is an operational definition: a Skillware Unit is the independently managed software identity wrapped around a Skill Artifact, and membership is decided by three necessary conditions—C1 behavioral primacy, C2 independent software identity, C3 agent-host execution—with Lifecycle Continuity evaluated as an orthogonal software-grade property. The definition anchors a chain of responsibilities from behavioral source, through skill artifact and unit, to agent host, runtime, execution trace, and task outcome, which keeps artifact specification separate from runtime interpretation and situated performance.","core_discovery":"The paper claims that persistent natural-language behavioral specifications create a managed software surface in agent systems. It defines Skillware as the software abstraction that extends software engineering to this surface by providing addressable artifacts, independent identity, documented or reconstructed activation relations, engineering structure, and a basis for identity-preserving change. Operationally, membership requires three necessary conditions applied to the same unit: behavioral primacy (the skill source organizes the reusable task contract), independent software identity (a name, address, package, version, or provenance trackable apart from the agent host), and an agent-hos","pith_inferences":["If the ontology is adopted, natural-language behavioral specifications could eventually be published with semantic versioning and compatibility declarations, letting agents check whether a skill will work before activation.","The C1-C3 boundary suggests a practical diagnostic: many files in skill registries may fail C2 or C3, meaning the true population of Skillware units could be much smaller than file counts imply.","A testable extension would measure Lifecycle Continuity across a random sample of skill repositories; low continuity would indicate that most skills are still disposable prompts rather than managed software.","The ontology implies a clean division between artifact quality and runtime quality, which could change how failures are attributed in agent systems—ambiguous source versus model, context, or tool problems."],"forward_implications":["Agent skills can be treated as identifiable software units, so acquisition, installation, activation, update, and removal become managed operations rather than ad-hoc file copying.","The three conditions give an auditable boundary that distinguishes Skillware from prompts, tools, plugins, indexes, and session-bound configurations.","Lifecycle Continuity makes 'the same skill across versions' a measurable property, enabling versioned release, rollback, and maintenance tracking.","Established design patterns and engineering mechanisms can transfer to natural-language behavioral source, giving architects known structures for building and testing skill-based systems.","Engineering Consolidation and Identity-Preserving Evolution describe how a text-first skill can mature into a hybrid system with scripts, tests, events, and governed change while keeping one identity."],"fun_headline_variants":["Agent skills become software with Skillware","Skillware: persistent behavior as managed software","Three conditions certify skills as software","Ontology gives agent skills a software lifecycle","Skillware extends SE to persistent skills"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the sampled skill files and reviewed cases fairly represent the broader ecosystem; if the registry-heavy corpus and purposively selected cases are unrepresentative, or if other analysts would classify the boundaries differently, the claimed universality of the Skillware ontology is not established.","fun_headline_variants_meta":{"raw":{"variants":["Agent skills become software with Skillware","Skillware: persistent behavior as managed software","Three conditions certify skills as software","Ontology gives agent skills a software lifecycle","Skillware extends SE to persistent skills"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1350,"prompt_tokens":782,"completion_tokens":568,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":506}},"tokens_in":526,"tokens_out":568,"duration_ms":5153,"temperature":1.0,"reasoning_tokens":506,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T13:47:14.538946+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent raters apply C1-C3 to a random, revision-frozen sample of skill-bearing repositories, including ones with heavy executable support; if a substantial share fail C2 (no independently trackable identity) or C3 (no skills-compatible agent host activation path), or if raters disagree systematically on C1, then the claim that persistent behavioral artifacts form a recurring managed software surface would be refuted.","supporting_citations":[],"review_version":1}