Pith. sign in

REVIEW 3 major objections 5 minor 34 references

Studying Developer Perceptions on the Potential of CI Recommendation Systems

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper argues that no one has measured whether developers would use CI recommendation systems, and it designs the first survey to do so.

desk verdict A well-specified survey protocol for a real gap, but the power analysis leans on optimistic sampling assumptions and the ethics section contradicts itself. read the letter →

arxiv 2608.02682 v1 pith:3EFX6JBZ submitted 2026-08-03 cs.SE

classification cs.SE
keywords ContinuousIntegrationCIadoptionservicesrecommendationsystemdevelopersurveyempiricalsoftwareengineeringGitHubTechnologyAcceptanceModel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Continuous integration (CI) is now a default practice in open-source development, yet the paper argues that two basic questions remain unanswered: whether projects adopt CI and choose a CI service because of genuine project needs or because of social influence, and whether developers would actually use automated recommendation systems that suggest a CI service for a project. Earlier surveys reached only CI users, treated non-adopters as a small minority, or predated the GitHub Actions era, so no empirical evidence exists on these questions. This paper responds by proposing an exploratory survey of roughly 5,000 active GitHub developers, expecting 250–500 complete responses split roughly 40–60% between CI users and non-users. The study plan measures adoption motivations, beliefs about whether CI is universally necessary or context-dependent, barriers among non-adopters, and perceived usefulness, trust, and adoption likelihood for CI recommendation systems. If the data arrive as planned, the result would be the first empirical evidence on developer perceptions of CI recommendation systems and the first comparison of CI users and non-users on those perceptions.

What carries the argument

The load-bearing object is a two-part conditional-routing survey instrument (Part A for CI users, Part B for non-CI users) delivered through a web survey platform to a GitHub-sampled population of roughly 5,000 active developers. The routing question is what makes the comparison possible: it admits non-adopters, unlike prior CI surveys. The perceived-value items for usefulness, trust, and likelihood of use operationalize the Technology Acceptance Model for a hypothetical CI recommendation system. The planned quantitative machinery is Mann-Whitney U tests with effect sizes, Chi-square tests on categorical answers, and thematic analysis with inter-rater agreement measured by kappa for open-ended responses. The sample-size constraint ($n \geq 87$ per group) is what connects the expected 40–60% CI-user split to the stated statistical power.

What would settle it

Track response rates separately for invited developers who are CI users and those who are not; if fewer than $87$ complete non-CI responses come in while CI-user responses exceed that, the planned Mann-Whitney U comparisons between the two groups cannot be run at the stated power. The same check applies if the overall response rate falls below 5% or the non-CI share drops far below 40%.

Watch

Extended reading notes

Core claim

The central claim is that the research community currently lacks, and can obtain, empirical evidence on whether developers would welcome automated CI recommendation systems and on whether CI adoption is driven by need or social influence. Prior surveys sampled almost exclusively CI users and predated the current service landscape, leaving non-adopter perspectives and recommendation-system perceptions unmeasured. The paper's design targets about 5,000 active open-source developers, uses a routing question to split respondents into CI-user and non-CI groups, and plans subgroup comparisons powered by a sample-size calculation requiring at least $87$ usable responses per group. The authors state as an anticipated result that non-CI users may perceive higher value in recommendation systems than CI users, and that developers may view CI as context-dependent rather than universally necessary.

Load-bearing premise

The plan assumes that enough non-CI developers will respond to form a group of at least $87$, because if non-CI developers respond far less often than CI users, the promised user-versus-non-user comparison loses the statistical power the paper relies on.

Editorial extensions

If this is right

  • The survey's first dataset, if it reaches 250–500 responses, will be the first empirical basis for deciding whether CI recommendation tools should be built at all and which features they should offer.
  • A confirmed split between need-driven and socially influenced adoption would let tool designers weight service-selection criteria by what actually drives decisions instead of defaults.
  • If developers treat CI as context-dependent, a recommendation system should assess project suitability before suggesting CI, rather than recommending it for every project.
  • If non-CI users report higher perceived value than CI users, separate recommendation paths for the two groups would be justified.
  • Identified barriers among developers who believe CI is useful only in some contexts would give tool builders a concrete list of obstacles to remove.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not connect its three research questions in a moderation analysis; a natural extension would be to test whether developers who cite social influence as a motivation are also more likely to switch services or use multiple CI services.
  • Measuring perceived usefulness and trust at one point captures intention, not adoption; a follow-up field study where a recommendation tool is actually deployed would test whether stated willingness predicts use.
  • Because the sampling frame is restricted to active open-source developers with public emails, the results likely describe that population; comparing early and late respondents or adding an enterprise sample could test how far the evidence generalizes.
  • If usefulness ratings are high but trust ratings are low, the implied design fix is a system that explains its recommendations; the survey measures both constructs but does not test that link.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This manuscript proposes a large-scale survey-based study of approximately 5,000 active GitHub developers (with a target of 250–500 responses) to investigate three research questions: (RQ1) what drives CI adoption and service selection, distinguishing need-driven from socially influenced motivations; (RQ2) whether developers view CI as universally necessary or context-dependent and what barriers prevent adoption; and (RQ3) how developers perceive the usefulness, trustworthiness, and desired features of automated CI recommendation systems. The paper presents the study design, sampling strategy, survey instrument, quantitative and qualitative analysis plans, execution phases, ethical considerations, and threats to validity. It claims that, once executed, the study will provide the first empirical evidence on developers' perceptions of CI recommendation systems.

Significance. If the study were executed as described, it would address a genuine gap in the literature: prior large-scale CI surveys (Hilton et al. 2016, Pinto et al. 2018, Widder et al. 2019) focused predominantly on CI users and did not ask about automated CI recommendation systems. The explicit inclusion of non-CI users, the pre-specified analysis plan, the power analysis, and the grounding of RQ3 in the Technology Acceptance Model and trust are strengths. The paper also provides a detailed survey instrument and an execution plan that could serve as a useful template for replication. However, the contribution is entirely prospective: no data are collected or analyzed, so the claim that the paper 'will provide the first empirical evidence' is not yet a scientific finding but a promise. The significance of the contribution therefore rests entirely on the execution plan and the plausibility of its sampling and power assumptions.

major comments (3)
  1. [§3.2] The power analysis and sample-composition assumption are not conservative for the planned comparison between CI users and non-CI users. The sampling frame selects active committers in repositories with at least 10 stars, a population likely to skew toward developers with CI experience, and non-CI developers may be less inclined to respond to a CI-focused survey. If the realized split is 70/30 instead of the assumed 40–60% CI users, a minimum sample of 250 responses yields only 75 non-CI respondents, below the n=87 per group required for the Mann-Whitney U test at r=0.3 and 80% power. The proposed mitigation—'we may prioritize non-CI developers'—is under-specified: CI status is not observable from the GitHub API before the screening question, so targeted invitations would require a proxy (e.g., sampling repositories without CI configuration files), which is not described. The authors should provide a sensitivity analysis across plausible CI-user proportions, or a concrete mechanism for oversampling non-CI developers, or explicitly weaken the cross-group comparison claim.
  2. [§3.5 vs. §3.6] There is a direct internal contradiction about ethics approval status. Section 3.5, Phase 1 states 'we have prepared the full survey protocol ... and submitted it for ethics approval,' while Section 3.6 states 'An ethics application for this study has been approved by Trent University.' These statements cannot both be true. The paper must clarify the actual current status (submitted, pending, or approved), and if approval has been granted, provide the REB reference or approval number. This is a load-bearing issue because the manuscript explicitly invokes ethical compliance as part of its methodology.
  3. [Abstract and Conclusion] The central claim that this study 'will provide the first empirical evidence on developers' perceptions of such systems' is unsupported by the manuscript's content, which contains no empirical data or results. If the paper is intended as a study protocol or registered report, the title, abstract, and conclusion should be reframed to present the contribution as the design, instrument, and analysis plan, with the evidence-generation claim explicitly deferred to future work. If it is intended as a complete empirical study, the data and analysis are missing. As written, the claimed contribution is not commensurate with the actual content.
minor comments (5)
  1. [§2.1.3] The term 'codeql' should be capitalized as 'CodeQL' when referring to the tool.
  2. [§3.2] The sentence 'This 5,000 represents oursampling population' contains a missing space; it should read 'our sampling population.'
  3. [§3.4] The qualitative analysis plan states that 'the two co-authors will collaboratively code all responses' but also that 'we will assess inter-rater reliability using Cohen's kappa.' Collaborative coding typically does not yield meaningful inter-rater reliability; the authors should clarify whether coding is independent with disagreement resolution or consensus-based, and adjust the reliability assessment accordingly.
  4. [§3.2] The survey instrument is linked to a Google Docs URL; for archival purposes, a permanent repository or DOI would be more appropriate.
  5. [Table 1] The table distinguishes 'Hilton 2016' and 'Hilton 2017' but the reference list includes Hilton et al. 2016 and Hilton et al. 2017; please ensure the table entries correspond unambiguously to the cited works.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey design is prospective and self-contained; cited results, including self-citations, serve only as background motivation or sampling assumptions, not as inputs that determine the study's conclusions.

full rationale

This paper is a study proposal with no empirical results and no derivation chain, so there is no step in which a prediction is equivalent to its inputs by construction. The three research questions (RQ1–RQ3) and the survey instrument are defined independently of the anticipated outcomes. Self-citations such as Chopra and Ghaleb [6] and Ghaleb et al. [17] are used as prior empirical context and to calibrate expected response rates, but they do not force any particular survey result; the same is true for the external surveys by Hilton et al., Pinto et al., and Widder et al. The Section 3.2 assumption of a 40–60% CI-user split and a 5–10% response rate is a sample-size planning assumption, not a fitted parameter disguised as a prediction. The 'knowledge gap' claim that little is known about developer perceptions of CI recommendation systems is supported by comparison with prior surveys that did not ask those questions, and the paper's own prior work is not invoked as the sole or load-bearing justification for that gap. No fitted input is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via citation. The study is therefore self-contained against external benchmarks, and the appropriate circularity score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The paper produces a study design, not a derivation, so the ledger records the planning assumptions. The main free parameters are response-rate and adoption-rate targets borrowed from prior literature; they are not fitted to new data. The axioms are standard statistical methods and domain assumptions about GitHub sampling and self-reported CI status. No invented entities are introduced.

free parameters (3)
  • Planned response rate 5-10%
    Used to target 250-500 responses from 5,000 invitations; based on prior survey response rates (Hilton 9.8%, Pinto 14.4%, Ghaleb 3.5%).
  • Assumed CI-user proportion 40-60%
    Used to estimate subgroup sizes (100-150 CI users for n=250) for power; taken from prior adoption estimates.
  • Assumed effect size r=0.3
    Used in power analysis to justify n>=87 per group; Cohen's medium effect convention.
assumptions (4)
  • domain assumption GitHub REST API sampling of repos with >=10 stars and >=10 commits in the past 12 months selects active developers
    Section 3.2; follows Kalliamvakou et al. [25] guidelines for GitHub mining.
  • domain assumption Self-reported CI involvement is a valid dependent measure
    Section 3.7 External validity: the authors rely on self-reports rather than repository CI config detection.
  • standard math Mann-Whitney U and chi-square tests are appropriate for Likert and categorical comparisons
    Section 3.3 Quantitative Analysis.
  • standard math Thematic analysis with inter-rater agreement (Cohen's kappa >= 0.70) yields reliable qualitative findings
    Section 3.4 Qualitative Analysis.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Studying Developer Perceptions on the Potential of CI Recommendation Systems." pith.science (2026). https://pith.science/paper/3EFX6JBZ

@misc{pith2026260802682,
  author       = {Pith},
  title        = {Pith review of: Studying Developer Perceptions on the Potential of CI Recommendation Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3EFX6JBZ}},
  note         = {Machine review of arXiv:2608.02682}
}
read the original abstract

Continuous Integration (CI) is central to modern software development, yet developers often struggle to choose the most suitable CI service. Prior work has identified barriers to CI adoption but offers little empirical evidence on how developers select CI services or whether adoption decisions are driven by genuine project needs versus social influence. This paper presents an exploratory survey study addressing that gap. We aim to contact about 5,000 active GitHub developers, including both CI users and non-users. The study investigates: (1) what drives CI adoption and service selection, distinguishing need-driven from socially influenced motivations; (2) whether developers consider CI universally necessary or context-dependent and what barriers hinder adoption; and (3) developers' perceptions of automated CI recommendation systems. Our findings will inform researchers developing CI recommendation systems and practitioners aiming to streamline CI adoption in open-source projects.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

34 extracted references · 32 canonical work pages

  1. [1]

    2006 , url =

    Martin Fowler , title =. 2006 , url =

  2. [2]

    2010 , publisher=

    Continuous delivery: reliable software releases through build, test, and deployment automation , author=. 2010 , publisher=

  3. [3]

    Proceedings of the 31st IEEE/ACM international conference on automated software engineering , pages=

    Hilton, Michael and Tunnell, Timothy and Huang, Kai and Marinov, Darko and Dig, Danny , title=. Proceedings of the 31st IEEE/ACM international conference on automated software engineering , pages=

  4. [4]

    Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering , pages=

    Hilton, Michael and Nelson, Nicholas and Tunnell, Timothy and Marinov, Darko and Dig, Danny , title=. Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering , pages=

  5. [5]

    Work practices and challenges in continuous integration: A survey with

    Pinto, Gustavo and Castor, Fernando and Bonifacio, Rodrigo and Rebou. Work practices and challenges in continuous integration: A survey with. Software: Practice and Experience , volume=. 2018 , publisher=

  6. [6]

    A conceptual replication of continuous integration pain points in the context of

    Widder, David Gray and Hilton, Michael and K. A conceptual replication of continuous integration pain points in the context of. Proceedings of the 2019 27th acm joint meeting on european software engineering conference and symposium on the foundations of software engineering , pages=

  7. [7]

    2025 IEEE International Conference on Software Maintenance and Evolution , pages=

    Chopra, Nitika and Ghaleb, Taher A , title=. 2025 IEEE International Conference on Software Maintenance and Evolution , pages=. 2025 , organization=

  8. [8]

    IEEE Transactions on Software Engineering , volume=

    Elazhary, Omar and Werner, Colin and Li, Ze Shi and Lowlind, Derek and Ernst, Neil A and Storey, Margaret-Anne , title=. IEEE Transactions on Software Engineering , volume=. 2021 , publisher=

Show all 34 references
  1. [9]

    Empirical Software Engineering , volume=

    Ghaleb, Taher Ahmed and Da Costa, Daniel Alencar and Zou, Ying , title=. Empirical Software Engineering , volume=. 2019 , publisher=

  2. [10]

    IEEE Transactions on Software Engineering , volume=

    Ghaleb, Taher A and Hassan, Safwat and Zou, Ying , title=. IEEE Transactions on Software Engineering , volume=. 2022 , publisher=

  3. [11]

    IEEE Transactions on Software Engineering , volume=

    Gallaba, Keheliya and McIntosh, Shane , title=. IEEE Transactions on Software Engineering , volume=. 2018 , publisher=

  4. [12]

    arXiv preprint arXiv:2507.20402 , year=

    Hossain, Md Nazmul and Ghaleb, Taher A , title=. arXiv preprint arXiv:2507.20402 , year=

  5. [13]

    2024 IEEE International Conference on Source Code Analysis and Manipulation , pages=

    Valenzuela-Toledo, Pablo and Bergel, Alexandre and Kehrer, Timo and Nierstrasz, Oscar , title=. 2024 IEEE International Conference on Source Code Analysis and Manipulation , pages=. 2024 , organization=

  6. [14]

    Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , pages=

    Bouzenia, Islem and Pradel, Michael , title=. Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , pages=

  7. [15]

    2022 IEEE International Conference on Software Analysis, Evolution and Reengineering , pages=

    Golzadeh, Mehdi and Decan, Alexandre and Mens, Tom , title=. 2022 IEEE International Conference on Software Analysis, Evolution and Reengineering , pages=. 2022 , organization=

  8. [16]

    2022 IEEE International Conference on Software Maintenance and Evolution , pages=

    Decan, Alexandre and Mens, Tom and Mazrae, Pooya Rostami and Golzadeh, Mehdi , title=. 2022 IEEE International Conference on Software Maintenance and Evolution , pages=. 2022 , organization=

  9. [17]

    Qualitative research in psychology , volume=

    Braun, Virginia and Clarke, Victoria , title=. Qualitative research in psychology , volume=. 2006 , publisher=

  10. [18]

    MIS quarterly , volume=

    Davis, Fred D , title=. MIS quarterly , volume=. 1989 , publisher=

  11. [19]

    Management science , volume=

    Venkatesh, Viswanath and Davis, Fred D , title=. Management science , volume=. 2000 , publisher=

  12. [20]

    Proceedings of the 11th working conference on mining software repositories , pages=

    Kalliamvakou, Eirini and Gousios, Georgios and Blincoe, Kelly and Singer, Leif and German, Daniel M and Damian, Daniela , title=. Proceedings of the 11th working conference on mining software repositories , pages=

  13. [21]

    Experimentation in software engineering , volume=

    Wohlin, Claes and Runeson, Per and H. Experimentation in software engineering , volume=. 2012 , publisher=

  14. [22]

    The annals of mathematical statistics , pages=

    Mann, Henry B and Whitney, Donald R , title=. The annals of mathematical statistics , pages=. 1947 , publisher=

  15. [23]

    2005 , publisher=

    Grissom, Robert J and Kim, John J , title=. 2005 , publisher=

  16. [24]

    arXiv preprint arXiv:2507.18062 , year=

    Abrokwah, Edward and Ghaleb, Taher A , title=. arXiv preprint arXiv:2507.18062 , year=

  17. [25]

    International Conference on Evaluation and Assessment in Software Engineering , year=

    Abrokwah, Edward and Ghaleb, Taher A , title=. International Conference on Evaluation and Assessment in Software Engineering , year=

  18. [26]

    IEEE Access , volume=

    Shahin, Mojtaba and Babar, Muhammad Ali and Zhu, Liming , title=. IEEE Access , volume=. 2017 , publisher=

  19. [27]

    Tri-Council Policy Statement: Ethical Conduct for Research Involving Humans -- TCPS 2 (2018) , year=

  20. [28]

    Indianapolis, Indiana , volume=

    Dillman, Don A and Smyth, Jolene D and Christian, Leah Melani , title=. Indianapolis, Indiana , volume=

  21. [29]

    Empirical Software Engineering , volume=

    Rostami Mazrae, Pooya and Mens, Tom and Golzadeh, Mehdi and Decan, Alexandre , title=. Empirical Software Engineering , volume=. 2023 , publisher=

  22. [30]

    ACM Transactions on Software Engineering and Methodology , volume=

    Ghaleb, Taher A and Abduljalil, Osamah and Hassan, Safwat , title=. ACM Transactions on Software Engineering and Methodology , volume=. 2026 , publisher=

  23. [31]

    1988 , publisher =

    Jacob Cohen , title =. 1988 , publisher =

  24. [32]

    Alaini and Taher A

    Osamah H. Alaini and Taher A. Ghaleb , title =. Companion Proceedings of the 34th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering , year =

  25. [33]

    Proceedings of the 23rd International Mining Software Repositories Conference , year=

    Ghaleb, Taher A , title=. Proceedings of the 23rd International Mining Software Repositories Conference , year=

  26. [34]

    International Conference on Evaluation and Assessment in Software Engineering , year=

    Ghaleb, Taher A and da Costa, Daniel Alencar and Zou, Ying , title=. International Conference on Evaluation and Assessment in Software Engineering , year=

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.