{"id":"60ec59e6-97c1-42c4-adea-dc0377901248","arxiv_id":"2506.08055","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper maps 62 existing tools, several proposed frameworks, and recurring security challenges across 66 reviewed papers in cloud-based CI/CD.","lead":"A systematic literature review of 66 papers maps the tools, proposed solutions, and challenges for securing CI/CD pipelines that run in the cloud. It is a compact starting point for practitioners and researchers trying to see the landscape of cloud CI/CD security.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Results tables cite sources outside the declared 66-paper corpus, and the corpus itself is never listed; the tool/challenge map may not be a systematic synthesis.","rationale":"The reader's verdict is CONDITIONAL, and I concur that the central risk is the unverifiable sample of 66 studies. My concern is aligned but more specific: the paper's own result tables appear to cite works that cannot be among the included studies, which would mean the RQ1/RQ2 map is not actually a synthesis of the declared corpus. Example: Table 2's 'Signature-based, Anomaly-based' entry lists references absent from the bibliography; also, including framework papers like STRIDE and Asylo under an empirical-research inclusion criterion suggests the criterion was not applied consistently. The paper has real strengths: it follows a recognizable Kitchenham-based SLR process, reports demographic data, and the 62-tool count in Table 2 is internally consistent. The counts in Tables 2 and 3 are internally coherent (62+5 tools, 8+12 approaches, matching row counts), which indicates care in the result tables themselves. The missing study list and the untraceable citations are the load-bearing vulnerability because they prevent any independent check of the central descriptive claims. A reviewer can neither confirm that the findings are comprehensive nor that they are drawn from the stated 66 papers. My recommended action is to keep the reader's CONDITIONAL verdict and specify the condition: the authors should supply the full study list, a PRISMA flow diagram with exclusion counts, and a traceability matrix from each table entry to included studies. Only after that check can the review be evaluated. This does not change the reader's verdict; it sharpens the condition.","tokens_in":15089,"tokens_out":8442,"duration_ms":91712,"concrete_test":"Obtain from the authors the complete list of the 66 included primary studies and a PRISMA-style flow diagram with exclusion reasons. Then build the union of all citations appearing in Tables 2 and 3 and check each citation against (a) that list and (b) the reference list. For each citation not in the 66, examine whether the cited paper is actually used elsewhere in the text as a primary study or was imported from outside the systematic selection. Re-derive the tool and approach counts using only studies on the 66-paper list; if any table entries are not backed by an included study, the corresponding RQ1/RQ2 finding should be removed or re-run. If the corrected counts differ from the reported '62 tools, eight approaches' and 'five tools, twelve approaches,' the central map changes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that a systematic review of 66 selected papers yields a map of CI/CD security tools, approaches, challenges, and research gaps. For that claim to hold, the reported findings must actually derive from the 66 included primary studies. Two problems undermine this. First, nowhere does the paper list the 66 included papers; there is no PRISMA-style enumeration, exclusion log, or quality assessment, so the sample is not auditable. Second, and more specifically, the results tables contain citations that are not traceable to the selected corpus. For example, Table 2's row for 'Signature-based, Anomaly-based' cites 'Jyothsna et al., 2011; Kumar and Sangwan, 2012,' neither of which appears in the reference list; other rows cite conceptual or framework papers (e.g., STRIDE, Asylo) despite the inclusion criterion in Section 3.4 requiring 'Empirical research (Kitchenham et al., 2022b).' If the tool and approach inventories in Tables 2 and 3 were populated from background knowledge or reference chasing rather than from the 66 vetted studies, then the RQ1 and RQ2 results are not a systematic synthesis of the selected literature. Section 2.2's research gaps are likewise drawn from prior reviews, not demonstrably derived from the included studies. This is not a cosmetic reporting omission: without traceability, the descriptive map and gap claims are unsupported by the declared method.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a systematic literature review (SLR) on the security of CI/CD pipelines in cloud computing. The authors define three research questions: existing tools and methods (RQ1), proposed solutions (RQ2), and challenges (RQ3). They report a search across six digital libraries, 4,889 initial records, 573 after screening, and 66 included primary studies. The results comprise tables of existing and proposed tools/approaches and a narrative of challenges, followed by discussion, validity threats, conclusions, and future work. The core claim is that the 66 selected studies provide a comprehensive map of the field and reveal research gaps in the security of cloud-based CI/CD pipelines.","tokens_in":15304,"tokens_out":5076,"duration_ms":49355,"significance":"If the synthesis were fully traceable, this SLR would fill a useful niche by focusing specifically on security of CI/CD in the cloud, complementing broader prior surveys such as Shahin et al. and Rajapakse et al. The paper follows an established SLR structure: PICOC-framed RQs, automated search over six libraries, snowballing, and demographic analysis. The demographic breakdown (Section 4) and the identification of specific challenges such as image manipulation, unauthorized access, and weak authentication are potentially useful. However, the contribution is currently unverifiable because the included-study list, exact search string, and quality assessment are not reported. The stated mapping of tools and challenges therefore cannot be audited, which limits the practical value of the review as a systematic secondary study.","major_comments":[{"comment":"The paper states 'After thoroughly reviewing the full articles, 66 were included in our final selection' but never lists these 66 papers or provides a PRISMA-style flow diagram with counts at each stage. This is load-bearing because RQ1-RQ3 results are aggregates over this sample; without the list, the reader cannot verify that the sample matches the inclusion criteria or that the tables and narrative in Section 4 indeed derive from the sample. Please add a supplementary list of included studies with unique identifiers and use those identifiers consistently in all results tables and in Section 4.3 challenges.","section":"Section 3.5, Step 3"},{"comment":"Several citations in the results tables are not traceable to the declared corpus or the reference list. For example, Table 2's 'Signature-based, Anomaly-based' row cites 'Jyothsna et al., 2011; Kumar and Sangwan, 2012,' neither of which appears in the References. The 'STRIDE' row cites Davis et al. (2022), which per the reference list is a paper on a first offering of a software engineering course, not an empirical study of secure CI/CD. These entries indicate that some table rows were not sourced from the 66 vetted primary studies, undermining the claim that the tool map is a systematic synthesis. Provide a row-by-row mapping to included studies; rows without such a mapping should be removed or moved to a clearly labeled background section.","section":"Tables 2 and 3"},{"comment":"The search string is only shown in an image (Figure 1) and is not provided as text. Reproducibility is a core requirement of an SLR; the reader cannot see the exact terms, Boolean operators, field restrictions, or date filters. Include the full search string for at least one digital library, the date of the search, and the number of hits per library.","section":"Section 3.2 and Figure 1"},{"comment":"No quality assessment of the 66 primary studies is reported, even though the method section cites Kitchenham et al. (2022a). Section 6 mentions variability in study design and quality but does not describe an assessment instrument, how quality was scored, or how it influenced the synthesis. Without this, the challenge synthesis in Section 4.3 gives equal weight to all sources. Add a quality checklist, report scores or a summary, and state how quality was used in the synthesis.","section":"Section 6"},{"comment":"The 'Research Gaps' subsection is framed as an outcome of this review, but the paragraphs cite prior literature (e.g., Garg and Stavik 2019; Rafi et al. 2022; Decan et al. 2022) rather than being demonstrably derived from the 66 included studies. It is unclear whether these gaps emerged from the included studies or from the authors' broader reading. Clarify the provenance of these gaps; if they come from related work, label them accordingly and show how the current SLR extends them.","section":"Section 2.2"}],"minor_comments":[{"comment":"The sentence '66 met our selection criteria (see Section 3.3)' points to the wrong section; the selection criteria are defined in Section 3.4, not Section 3.3.","section":"Section 1"},{"comment":"The numbers are confusing: the text says 573 papers remained after screening, of which '482 directly met our criteria, and an additional 91 were found using the Snowballing method.' Since 482 + 91 = 573, clarify whether the 573 already includes the snowballed papers, and provide exclusion counts for each reason.","section":"Section 3.5, Step 2"},{"comment":"The 'GitHub Actions (GHA)' row cites 'Tu et al., 2021,' but the reference list entry for Tu et al. (2021) is about Open vSwitch dataplane performance, which appears unrelated to GitHub Actions. Verify or replace this citation.","section":"Table 2"},{"comment":"There are inconsistent author names between text and references: 'Garg and Stavik, 2019' vs. 'Garg and Garg, 2019'; 'Brandy et al., 2020' vs. 'Brady et al., 2020'; and the reference list contains duplicate entries for Shahin et al. (2023). Standardize these.","section":"References"},{"comment":"The bullet 'Trust developers' cites Shahin et al. (2017b), a prior survey; if this is a finding from the included studies, provide the corresponding included-study ID, otherwise clarify that it is background advice.","section":"Section 4.2"},{"comment":"The study-selection figure is labeled 'Steps of Study Selection for SLR' but contains no numbers at each step. Replace or augment it with a flow diagram showing the counts for retrieval, screening, eligibility, and inclusion, as is standard in SLR reporting.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope as a secondary study, but the reporting gaps are substantial: no list of included studies, no exact search strings, no quality assessment, and several untraceable citations in the results tables. These issues are fixable with supplementary material and a thorough revision. I would be willing to review a revised version that addresses the traceability concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This one is useful as a map, but I'd hold the 'systematic' label until the authors open their box.\n\nWhat the paper does well: it picks a real niche—cloud-specific CI/CD security—that sits between Shahin et al. 2017 and Rajapakse et al. 2022, and it lists a practical range of tools (the 62-tool table includes Harbor, SonarQube, GitHub Actions, Trivy, etc.) and challenge categories (authorization, vulnerability assessment, third-party/OSS, supply chain). Practitioners starting in this area will find the tables genuinely useful, and the demographic data showing 60% of the 66 papers published 2021–2023 confirms this is a live topic. The process is recognizable: search strategy, digital libraries, snowballing, inclusion/exclusion, data extraction.\n\nThe soft spot is not the selection criteria; it's that the execution doesn't match the claim. There is no list of the 66 included studies, no PRISMA-style flow, no exclusion log, and no quality assessment, so nobody can audit the corpus. Worse, the results tables contain citations that are not in the reference list—'Jyothsna et al., 2011; Kumar and Sangwan, 2012' for signature/anomaly-based approaches don't appear anywhere in the references. Other rows cite conceptual works like STRIDE and Asylo even though the inclusion criteria require empirical research. That tells me the tables were partially filled from background knowledge or reference chasing, not strictly from the 66 vetted papers. If so, RQ1 and RQ2 are not a systematic synthesis; they're an informed newsletter. The same holds for Section 2.2's research gaps, which feel assembled from prior reviews rather than derived from the corpus.\n\nNone of this is fatal to the paper's usefulness as a first stop for someone entering the field. But it is fatal to its current claim to be a systematic literature review. The fix is straightforward: supply an appendix with the full list of 66 papers, a flow diagram, a mapping from each table row to an included study, and a quality assessment. If the authors can't do that, they should reframe the paper as a narrative survey.\n\nI'd send this to peer review—the topic matters, and a revised version could meet SLR standards. But the reviewer should demand the supplementary material before any acceptance. Don't cite it in its current form.","headline":"A useful first-stop map of CI/CD security tools and challenges, but the 'systematic' claim is not yet supported because the corpus is unauditable.","tokens_in":15868,"tokens_out":3856,"would_cite":false,"duration_ms":44085,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This systematic literature review claims that 66 selected papers provide a working map of tools, proposed solutions, and challenges for securing CI/CD pipelines in cloud environments.","keywords":["Continuous Integration","Continuous Deployment","CI/CD","Cloud Security","Systematic Literature Review","DevSecOps","Software Supply Chain","Container Security"],"falsifier":"A replication of the reported selection funnel, from 4,889 initial articles to 573 screened studies to 66 included papers, should reproduce the same 66 papers; if the included-paper list, which the paper does not publish, differs materially, or if the counts of 62 tools and eight approaches cannot be traced to the included papers, the map and gap claims are not stable.","tokens_in":14852,"feed_emoji":"🔐","tokens_out":9869,"duration_ms":100226,"temperature":0.7,"pith_summary":"This paper claims that a systematic review of 66 papers can map the current state of secure continuous integration and continuous delivery/deployment (CI/CD) in the cloud. Starting from 4,889 search results and filtering to 573 then 66 studies, it catalogues 62 tools and eight approaches used in practice, plus five proposed tools and twelve proposed frameworks. The review identifies recurring security problems, including image manipulation, unauthorised access, weak authentication, third-party and open-source dependency risks, and supply-chain attacks such as Log4j and SolarWinds, and argues these reveal research gaps. A reader should care because the map gives practitioners a concrete starting point for tooling choices and gives researchers a gap list to target.","feed_headline":"66 studies map security tools and gaps in cloud deployment pipelines","feed_subtitle":"The review catalogs 62 tools, eight approaches, and the unresolved security challenges of cloud deployment pipelines.","key_machinery":"The carrying mechanism is the systematic literature review protocol: research questions framed with PICOC (Population, Intervention, Comparison, Outcomes, Context), a search string run across six digital libraries, snowballing that added 91 candidate papers, and a three-step selection funnel from 4,889 initial articles to 573 screened studies to 66 included papers. This funnel defines the evidence base, and every catalog count, challenge category, and gap statement in the paper is an aggregation of what those 66 papers report.","core_discovery":"On its own terms, the paper's central finding is that the literature on CI/CD security in the cloud is rich in tools but thin on verified integration: the review compiled 62 tools and eight approaches or frameworks for existing practice, plus five proposed tools and twelve proposed approaches or frameworks. Challenges such as image manipulation, unauthorised access, weak authentication, disconnection between security tools and IDEs, third-party and open-source dependency risks, and supply-chain attacks such as Log4j, SolarWinds, and CodeCov recur across the selected studies. These recurring challenges, the paper argues, reveal research gaps in how tools and practices address security in CI/CD pipelines, justifying further study.","pith_inferences":["Because 40 of the 66 included papers were published in 2021 to 2023, the field is young, and the tool and challenge map may need to be re-run within a few years as GitHub Actions and supply-chain security mature.","The review's tool list is heavily Docker- and container-centric; a follow-up review targeting serverless, Kubernetes-native, or multi-cloud deployment models could find a different challenge profile.","The paper counts tools but does not compare their effectiveness; a natural testable extension would be to run a fixed vulnerability suite through the most-cited scanners to see which reported gaps actually close."],"forward_implications":["A practitioner can use the catalog of 62 tools and eight approaches as a starting checklist for container scanning, static and dynamic analysis, monitoring, and DevSecOps in cloud CI/CD.","The challenge categories of installation and updating, practitioner and developer issues, organisational issues, and third-party and open-source tool difficulties show where current tooling is least reliable.","GitHub Actions and low-code platforms emerge as specific weak points that need targeted hardening and further study.","The paper's planned next steps, topic modeling of fragmented security text and a blockchain-based solution for container and deployment security, follow directly from the identified gaps."],"supporting_citations":[{"why":"Supplies the SLR guidelines the protocol follows.","marker":"Kitchenham et al. (2022a)"},{"why":"Guides the search strategy and study identification.","marker":"Zhang et al. (2011)"},{"why":"Defines the snowballing approach that added 91 candidate papers.","marker":"Wohlin (2014)"},{"why":"Underpins the full-text, English-only inclusion criteria.","marker":"Brereton et al. (2007)"},{"why":"Earlier review of CI/CD adoption that this review extends by focusing on security in the cloud.","marker":"Shahin et al. (2017a)"},{"why":"Provides the validity-threat framework used in Section 6.","marker":"Runeson and Höst (2009)"}],"fun_headline_variants":["62 tools mapped, security gaps found in 66 CI/CD studies","Cloud CI/CD security: 66 studies, 62 tools, open gaps","Supply-chain risks and weak auth top cloud CI/CD gaps","Security tools outpacing verified practice in cloud CI/CD","66 papers, 62 tools: cloud CI/CD security still gapped"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 66 papers that survived the selection process are a complete and unbiased representation of the literature on CI/CD security in the cloud, since no exclusion log or quality assessment is reported.","fun_headline_variants_meta":{"raw":{"variants":["62 tools mapped, security gaps found in 66 CI/CD studies","Cloud CI/CD security: 66 studies, 62 tools, open gaps","Supply-chain risks and weak auth top cloud CI/CD gaps","Security tools outpacing verified practice in cloud CI/CD","66 papers, 62 tools: cloud CI/CD security still gapped"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000556,"raw_usage":{"total_tokens":2613,"prompt_tokens":875,"completion_tokens":1738,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":1649}},"tokens_in":491,"tokens_out":1738,"duration_ms":11129,"temperature":1.0,"reasoning_tokens":1649,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:34:21.519460+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication of the reported selection funnel, from 4,889 initial articles to 573 screened studies to 66 included papers, should reproduce the same 66 papers; if the included-paper list, which the paper does not publish, differs materially, or if the counts of 62 tools and eight approaches cannot be traced to the included papers, the map and gap claims are not stable.","supporting_citations":[{"cited_title":"A., & Tell, P","cited_arxiv_id":null,"evidence_quote":"Guides the search strategy and study identification."}],"review_version":1}