{"id":"1476891a-7327-4f56-a6ee-556f4e91fe56","arxiv_id":"2507.05100","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The first systematic taxonomy of Everything as Code, classifying 25 practices into six functional layers and a conceptual model, built from a multivocal literature review and expert validation.","lead":"This paper reviews 128 sources on 'everything as code' and organizes 25 practices, from infrastructure to security, into a six-layer taxonomy with a conceptual model of how they interact. It gives the software industry a common vocabulary and a starting map for adopting as-code practices across the delivery lifecycle.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 25-practice scope and established/emerging split rest on a corpus whose grey-literature arm was deliberately vendor-prioritized; if a broader sample shifts frequencies or surfaces additional practices, 'first comprehensive' weakens.","rationale":"The reader's weakest-assumption identification is on target: corpus representativeness is the load-bearing premise. I partially agree with the reader, but I would sharpen the concern in two ways. First, the vendor-prioritized grey-literature search combined with iterative keyword expansion is not merely a sampling bias but a self-reinforcing feedback loop: practices already present in vendor sources are more likely to be found again and then used as keywords for further searches. Second, the post hoc 20%-40% threshold and the unreported exact frequencies make the established/emerging classification unauditable from the paper alone. The paper is transparent about many limitations (Section VI), and the supplementary resources are a real strength. The expert validation is small and vendor-aligned, but since the taxonomy is primarily literature-derived, the expert panel is a refinement step rather than the main evidence. The same corpus concern that justifies a conditional verdict does not, without a concrete replication showing instability, justify rejection. Therefore I do not change the reader's verdict: CONDITIONAL remains appropriate, with the condition that the corpus-dependence of the 25-practice enumeration and the established/emerging split be empirically demonstrated or acknowledged more prominently.","tokens_in":14351,"tokens_out":4635,"duration_ms":56581,"concrete_test":"Using the supplementary source list (Section VII), recompute the per-practice literature frequencies and tool counts after (a) removing all grey-literature sources from vendor domains (e.g., hashicorp.com, redhat.com, aws.amazon.com) and (b) adding a balanced sample of non-vendor practitioner discourse (conference talks, Stack Overflow, GitHub discussions, non-English blogs) that satisfies the stated inclusion criteria. If any practice crosses the 20% boundary, or if a new as-code practice such as Law as Code or Contracts as Code qualifies, the taxonomy's scope and established/emerging split are corpus artifacts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('first comprehensive taxonomy and conceptual model') depends on two premises: (1) the 25 practices are the complete set of EaC practices in software delivery, and (2) the established/emerging classification reflects industry awareness. Both are derived from a corpus whose grey-literature arm (Section III-A) 'prioritized sources from leading cloud providers... and DevOps vendors' and used iterative keyword expansion. This creates a feedback loop: vendor-marketed practices (e.g., Policy as Code, Security as Code) are more likely to be found and recursively searched, inflating their 'literature frequency', while less marketed practices are under-indexed. Table IV lists Privacy as Code with 'N/A' tooling, yet Table V records expert feedback that Law as Code, Contracts as Code, and Management as Code were excluded because they 'lack sufficient literature support'. That exclusion is an artifact of corpus construction, not independent evidence of irrelevance. The 20%-40% frequency threshold and 'more than three tools' cut are post hoc, and exact per-practice frequencies are not reported in the text, so the classification cannot be audited from the paper alone. If a non-vendor-prioritized sample changes frequencies or surfaces additional practices, the '25 distinct practices' enumeration and the established/emerging classification are corpus-dependent rather than robust properties of the field.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a multivocal literature review (MLR) of Everything as Code (EaC), synthesizing 42 peer-reviewed papers and 86 grey literature documents into a taxonomy of 25 EaC practices and a conceptual model of their relationships. The taxonomy has two dimensions: industry awareness and tooling support, which separates established from emerging practices, and functionality and application, which organizes practices into six layers. The conceptual model focuses on the six established practices and is validated, together with the taxonomy, through expert review with three practitioners from a single cloud provider. The authors claim this is the first comprehensive taxonomy and conceptual model of EaC and provide supplementary materials including extracted data, validation checklists, and code examples.","tokens_in":14593,"tokens_out":5390,"duration_ms":64524,"significance":"If the taxonomy and conceptual model are robust, they would provide a useful structured vocabulary for a fragmented and rapidly growing area of software engineering practice. The paper has genuine strengths: it follows established MLR guidelines, uses a recognized taxonomy development method, documents data extraction and validation steps, provides supplementary resources, and explicitly acknowledges several limitations. The distinction between established and emerging practices and the proposed relationship between Policy as Code, Compliance as Code, and Security as Code could be a helpful baseline for practitioners and researchers. However, the central claims of '25 distinct practices' and of an accurate industry-awareness classification depend on corpus construction and threshold choices that are not sufficiently auditable from the paper as written.","major_comments":[{"comment":"The grey-literature arm of the corpus was explicitly constructed to prioritize sources from leading cloud providers, DevOps vendors, and technical blogging platforms, and the search was iteratively expanded from practices found in that same corpus. This creates a feedback loop in which vendor-marketed practices are more likely to be encountered, recursively searched, and hence counted as frequent, while less-marketed practices are under-indexed. The exclusion of Law as Code, Contracts as Code, and Management as Code on the ground that they 'lack sufficient literature support' (Table V) is therefore partly an artifact of the collection strategy rather than independent evidence of irrelevance. The paper should report per-practice literature frequencies and tool counts, and should demonstrate robustness by analyzing a non-vendor-prioritized sample or by explicitly bounding the claim to the vendor-oriented discourse it sampled.","section":"Section III-A and Table V"},{"comment":"The classification into Established and Emerging Practices depends on two thresholds: a literature frequency in the 'upper frequency band (20% - 40%)' and 'more than three documented tools'. The paper does not report the exact frequency of each of the 25 practices, so the reader cannot determine why, for example, IAM as Code is Emerging rather than Established, or why the 20% and 40% boundaries were chosen. The violin plot in Fig. 4 is not a substitute for a data table. Please include the complete frequency and tooling table and a sensitivity analysis showing how the classification changes under reasonable alternative thresholds.","section":"Section IV-A"},{"comment":"The validation stage involved only three engineers from a single cloud service provider, all with DevOps and cloud-native profiles. While the authors acknowledge this limitation in Section VI-D, the taxonomy's second dimension is explicitly called 'Industry Awareness and Tooling Support'; a three-person, single-organization panel cannot validate industry awareness across sectors. The claim that 'every element of the taxonomy and model had been explicitly ratified' should be toned down, or the panel should be expanded to include, for example, practitioners from enterprises without a cloud-vendor affiliation, plus researchers with relevant expertise.","section":"Section VI-A and Section VI-D"},{"comment":"The count of '25 distinct EaC practices' is internally inconsistent with the treatment of Storage as Code and Network as Code in Dimension 2. Section IV-B states that these practices 'were subsumed into Infrastructure as Code', yet they remain separate entries in Table IV and Fig. 6 and are counted among the 25. The paper should define whether the taxonomy enumerates leaf practices, umbrella practices, or both, and adjust the count and figures accordingly. This is load-bearing because the abstract promises '25 distinct EaC practices'.","section":"Section IV-B, Table IV, and Fig. 6"},{"comment":"The conceptual model is built only from the six established practices, as Section V states that it represents relationships between 'the established practices'. The other 19 identified practices are outside the model. Given that the title and abstract present the conceptual model as covering EaC as a whole, the model's scope should either be expanded to include emerging practices or the claims should be narrowed to a model of the established core. This distinction should be made explicit in the abstract and conclusion.","section":"Section V"}],"minor_comments":[{"comment":"The text contains a typo: 'Y AML' should be 'YAML'.","section":"Section V-B-2"},{"comment":"Several emerging practices are supported by a single grey-literature source (e.g., Storage as Code G57, IAM as Code G67, Privacy as Code G68). Please mark single-source practices as tentative or provide additional corroborating sources.","section":"Table IV"},{"comment":"The statement that the selected literature is 'available in Section VII' is unclear because Section VII refers to supplementary resources rather than listing all documents; please clarify what is in the paper versus the online repository.","section":"Section III-A"},{"comment":"The violin plot needs axis labels, a legend, and ideally a companion table of per-practice frequencies; the term 'trimodal distribution' cannot be inspected or verified from the figure as presented.","section":"Fig. 4"},{"comment":"Some references, particularly the book chapters [13]-[15], lack complete bibliographic details such as page numbers or DOIs; please make the reference list consistent.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's central contribution is plausible, but the 'first comprehensive' claim is broader than the evidence supports as currently presented. The authors should be asked either to qualify the claims or to provide the additional analysis, including per-practice frequencies and a robustness check of the classification. I did not access the GitHub repository; the per-practice frequencies and tool counts should appear in the paper itself so that the classification is auditable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a solid, honest systematic taxonomy paper for a field that genuinely lacks one. If you work on IaC/DevOps, you'll want it on the shelf. The 'first comprehensive' claim is defensible only within the boundaries of their corpus, but the paper itself mostly acknowledges that.\n\nWhat's new: they run a proper multivocal literature review (42 academic + 86 grey sources), apply Nickerson's taxonomy method, and end with 25 as-code practices in six functional layers plus a lifecycle conceptual model. Prior work covers subsets — IaC/CaC, or PaC/CoaC — nobody has assembled the full map. The supplementary GitHub materials (extracted data, validation checklists, JSON taxonomy, code examples) are real evidence and make the whole thing auditable. I take that seriously.\n\nThe soft spots are real but not load-bearing. The grey literature arm prioritized leading cloud providers and DevOps vendors, and the keyword expansion was iterative. So vendor-marketed practices like Policy as Code or Security as Code are more likely to show up in the corpus, and their frequency bands are partly an artifact of sampling. The 20%-40% established-practice threshold and 'more than three tools' cut were set post hoc, and per-practice frequencies aren't in the paper, so you can't fully audit the established/emerging split from the text alone. The expert validation was three engineers from one cloud provider — narrow, though they report the checklist results openly. And practices like Law as Code or Contracts as Code were excluded for 'lack of literature support,' which is circular given the corpus construction. That said, the authors flag all of this in Section VI. It's a starting point, not a finished ISO standard.\n\nBottom line: the taxonomy is corpus-dependent, but every taxonomy is. The contribution is a structured vocabulary and a testable baseline, not a law of nature. I would send this to referees — a serious SE venue can push them to report frequencies, broaden the grey lit, and soften the 'comprehensive' framing. I'd cite it in my own work if I were writing about EaC.","headline":"A solid, transparent taxonomy of Everything as Code that deserves refereeing; the 'first comprehensive' claim holds only within a vendor-influenced corpus the authors themselves acknowledge.","tokens_in":15127,"tokens_out":2062,"would_cite":true,"duration_ms":23399,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that Everything as Code has a definite structure: 25 distinct as-code practices organized into six functional layers and two maturity classes, presented as the first taxonomy and conceptual model of the field.","keywords":["Everything as Code","Infrastructure as Code","Policy as Code","Compliance as Code","Security as Code","multivocal literature review","taxonomy","conceptual model"],"falsifier":"A targeted check would be to run the same literature-frequency and tool-count analysis on a different corpus built from less vendor-centric sources, such as regional practitioner blogs, open-source documentation, and non-English technical writing. If any practice outside the 25 appears in the upper frequency band with more than three tools, or if one of the six established practices falls below those thresholds, the taxonomy's categories and boundaries fail their own test.","tokens_in":14146,"feed_emoji":"🧩","tokens_out":6018,"duration_ms":61173,"temperature":0.7,"pith_summary":"This paper claims that Everything as Code is not a vague umbrella term but a field with a definite structure. By systematically reviewing 42 peer-reviewed papers and 86 grey-literature documents, it identifies 25 distinct as-code practices and organizes them into six functional layers and two industry-awareness tiers. It presents the result as the first taxonomy and conceptual model of EaC, with a map of how the six established practices relate and overlap across the software delivery lifecycle. If the claim holds, practitioners get a shared vocabulary and a placement guide for tools and practices, and researchers get a baseline for a domain that currently lacks standards.","feed_headline":"Everything as Code: 25 practices, six layers, one map","feed_subtitle":"A systematic review shows how infrastructure, policy, security, compliance, and pipeline practices relate across the delivery lifecycle.","key_machinery":"The machinery that carries the argument is a two-dimensional taxonomy built through a standard taxonomy-development method with top-down and bottom-up passes. Dimension one counts literature frequency and tooling availability to divide the 25 practices into six established practices and nineteen emerging ones, using post-hoc thresholds: upper band of 20 percent to 40 percent frequency and more than three documented tools. Dimension two uses thematic analysis anchored to an established industry tooling landscape to define the six functional layers. The conceptual model translates extracted relationships into directional edges, with lifecycle alignment placing practices into DevOps stages; the single most load-bearing relationship is that Policy as Code validates Infrastructure as Code, which lets the model explain overlaps among Policy, Compliance, and Security as Code.","core_discovery":"The paper's central discovery is that Everything as Code decomposes into 25 discrete practices that can be sorted along two independent dimensions. The first dimension, industry awareness and tooling support, splits practices into Established and Emerging based on how often they appear in the literature and how many tools implement them; Infrastructure as Code is treated as an outlier at 71 percent frequency, and the remaining practices cluster into a 20 percent to 40 percent band versus a lower band. The second dimension, functionality and application, assigns each practice to one of six layers: Infrastructure Provisioning and Management, Platform and Orchestration, Application Design and Development, Data and Database, Security and Compliance, and Observability and Analysis. The conceptual model then adds directional relationships between the six established practices, such as Policy as Code validating Infrastructure as Code, and aligns them with DevOps lifecycle stages, while the paper gives explicit criteria for when a component belongs to Infrastructure as Code rather than Configuration as Code, using Kubernetes as the boundary case.","pith_inferences":["A fair stress test of the taxonomy would be to apply the same frequency and tooling thresholds to a fresh, independently collected corpus: if the established/emerging split moves, the current categories are corpus-specific rather than field-wide.","The machine-readable JSON version of the taxonomy points toward a living registry where new practices are classified by semi-automated means; the paper only proposes this as future work, but it is the natural next step.","Because the conceptual model covers only the six established practices, the nineteen emerging practices remain unconnected; extending the same directional-edge notation to them would complete the picture.","The paper's Kubernetes boundary example suggests a testable principle: as more components become 'infrastructure to applications,' the line between IaC and CaC will keep shifting, so the taxonomy may need versioned updates tied to technology shifts."],"forward_implications":["The taxonomy gives practitioners a checklist: any proposed as-code practice can be placed in one of the 25 slots or flagged as speculative until it gains literature and tooling support.","The established-versus-emerging split becomes an adoption roadmap: start with IaC, CaC, Pipeline as Code, Policy as Code, Security as Code, and Compliance as Code, which have tool ecosystems, before investing in emerging practices.","The IaC/CaC boundary criteria yield a concrete decision rule: configure a software component with IaC only when it is inseparable from infrastructure, foundational to applications, and part of a unified infrastructure unit; otherwise manage it as Configuration as Code.","The overlap analysis implies that one cloud resource may legitimately be scanned by security tools, validated by policy engines, and audited by compliance scripts in the same pipeline, each serving a different governance purpose.","The conceptual model's directional edges map directly to pipeline stages, so the model doubles as a design template for CI/CD automation."],"supporting_citations":[{"why":"Supplies the multivocal literature review guidelines that define how academic and grey literature are selected and synthesized.","marker":"[21]"},{"why":"Provides the taxonomy development method whose top-down and bottom-up directions determine the two dimensions.","marker":"[23]"},{"why":"Updates the taxonomy method with revised design guidance and ending conditions used to finalize the 25-practice set.","marker":"[24]"},{"why":"The closest prior work identifying IaC and CaC as EaC subsets, which this paper extends to the full practice set.","marker":"[3]"},{"why":"Frames Compliance as Code with Policy as Code downstream, one of the competing definitions the paper reconciles.","marker":"[10]"},{"why":"Defines Compliance as Code through OSCAL, supplying the academic anchor for that practice.","marker":"[18]"},{"why":"The industry landscape whose layers anchor the second dimension's six functional categories.","marker":"[26]"},{"why":"Documents Terraform provisioners, supporting the paper's argument about where IaC and CaC boundaries blur.","marker":"[27]"}],"fun_headline_variants":["Everything as Code: 25 practices mapped across six layers","Study maps Everything as Code into 25 practices, six layers","25 practices, six layers: a new taxonomy for Everything as Code","Everything as Code taxonomy: 25 practices, six layers, one model","How Everything as Code works: 25 practices and six layers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the selected corpus of 42 academic papers and 86 grey-literature documents, gathered with the stated inclusion and exclusion criteria and with searches that favored leading cloud and DevOps vendors, fairly represents the whole Everything-as-Code conversation; if that corpus is skewed, the 25-practice scope and the established/emerging split inherit the skew.","fun_headline_variants_meta":{"raw":{"variants":["Everything as Code: 25 practices mapped across six layers","Study maps Everything as Code into 25 practices, six layers","25 practices, six layers: a new taxonomy for Everything as Code","Everything as Code taxonomy: 25 practices, six layers, one model","How Everything as Code works: 25 practices and six layers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1415,"prompt_tokens":957,"completion_tokens":458,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":370}},"tokens_in":573,"tokens_out":458,"duration_ms":4565,"temperature":1.0,"reasoning_tokens":370,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:32:04.475208+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A targeted check would be to run the same literature-frequency and tool-count analysis on a different corpus built from less vendor-centric sources, such as regional practitioner blogs, open-source documentation, and non-English technical writing. If any practice outside the 25 appears in the upper frequency band with more than three tools, or if one of the six established practices falls below those thresholds, the taxonomy's categories and boundaries fail their own test.","supporting_citations":[{"cited_title":"Garousi, M","cited_arxiv_id":null,"evidence_quote":"Supplies the multivocal literature review guidelines that define how academic and grey literature are selected and synthesized."},{"cited_title":"An Update for Taxonomy Designers,","cited_arxiv_id":null,"evidence_quote":"Updates the taxonomy method with revised design guidance and ending conditions used to finalize the 25-practice set."},{"cited_title":"Toward Multiconcern Software Development With Everything as Code,","cited_arxiv_id":null,"evidence_quote":"The closest prior work identifying IaC and CaC as EaC subsets, which this paper extends to the full practice set."},{"cited_title":"Compliance-as-Code for Cybersecurity Automation in Hybrid Cloud,","cited_arxiv_id":null,"evidence_quote":"Defines Compliance as Code through OSCAL, supplying the academic anchor for that practice."},{"cited_title":"CNCF Landscape","cited_arxiv_id":null,"evidence_quote":"The industry landscape whose layers anchor the second dimension's six functional categories."},{"cited_title":"Provisioners, Terraform, HashiCorp Developer","cited_arxiv_id":null,"evidence_quote":"Documents Terraform provisioners, supporting the paper's argument about where IaC and CaC boundaries blur."}],"review_version":1}