Pith. sign in

REVIEW 5 major objections 4 minor 1 cited by

Are We on the Same Page? Examining Developer Perception Alignment in Open Source Code Reviews

T0 review · 5 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Open-source code-review friction comes less from bias than from contributors and maintainers valuing different things—maintainers want project fit, contributors lead with novelty.

desk verdict First survey comparison of contributor/maintainer review perceptions has a credible core, but the 'misperceived bias' claim rests on a definitional switch the authors don't acknowledge. read the letter →

arxiv 2504.18407 v1 pith:LWKFDZSL submitted 2025-04-25 cs.SE

classification cs.SE
keywords codereviewopensourcesoftwareperceptionalignmentbiascontributorsandmaintainersfamiliaritycontributionguidelinesmixedmethods
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that contributors and maintainers in open-source code review largely share one picture of what reviews should do, but systematically differ in emphasis, and that a large share of what developers call bias is really those emphasis gaps rather than discrimination. The evidence is a survey of 289 developers from 81 GitHub projects, interviews with 23 of them, and repository metadata; responses to identical questions are compared by role. Maintainers stress alignment with project goals, while contributors overvalue novelty and under-explain rationale. Among participants who reported bias, 65.40% described favoritism toward familiar contributors, while 37.02% described what the authors classify as differences in technical approach—experiences that do not meet the paper's working definition of bias. The paper draws practical consequences: better contribution guidelines, pre-review automation, and anonymized review should reduce both real bias and miscommunication.

What carries the argument

The mechanism is a paired perception-alignment study: separate role-specific surveys ask contributors and maintainers the same open-ended questions about review objectives, acceptance factors, challenges, and improvements; open coding turns free text into theme percentages; Spearman correlation measures overall alignment; and chi-square post-hoc tests locate which themes deviate significantly. The same machinery is applied to bias: respondents who said they witnessed bias were asked to describe it, and the authors classified descriptions into five categories (familiarity bias, approach difference, language challenges, misunderstanding, other). The categories do the paper's work—they separate signal (true bias) from noise (perceived bias caused by mismatched expectations), and the tabulated percentages become the paper's evidence that a large portion of reported bias is misattribution rather than discrimination.

What would settle it

Have two independent teams, blind to the paper's categories, re-code the same open-ended bias descriptions; if coders cannot agree, or if the same incident is placed in both 'familiarity bias' and 'approach difference,' the 37.02% misattribution figure has no stable referent. A second check: in follow-up interviews, ask participants whether they would still call the incident unfair once the reviewer's preference is explained as a technical style preference—if they do, the paper's signal/noise split does not capture how bias is experienced.

Watch

Extended reading notes

Core claim

The central claim is that perceived bias in open-source code review is a mix of signal and noise. The signal is real favoritism toward familiar contributors—named familiarity bias and reported by 65.40% of participants who noticed bias—which disproportionately hurts newer and underrepresented contributors. The noise consists of experiences the authors code as approach differences (37.02%), where a reviewer pushes for their preferred coding style, design, or technical solution; these are not unfair treatment under the definition the authors adopt, yet participants experience and describe them as bias. On the goals of review, maintainers and contributors agree on correctness, quality, standards, and documentation but diverge in emphasis: 31.37% of maintainers list alignment with project goals as an objective versus 13.37% of contributors, while 11.23% of contributors list novelty as a factor in pull-request acceptance versus 1.96% of maintainers, and only 3.74% of contributors mention rationale versus 12.75% of maintainers. The paper presents these role-based gaps in expectations as a cause of friction and disengagement that can be misread as bias.

Load-bearing premise

The argument depends on the assumption that independent coders can consistently tell a real bias experience apart from a mere difference in coding style or design preference, and that participants' word 'bias' means the same thing the coders' formal definition assumes; the paper states that coding disagreements were resolved by discussion rather than measured, so this classification is the step most worth checking.

Editorial extensions

If this is right

  • Writing project goals and the need for rationale directly into contribution guidelines should close the largest measurable gap between maintainer expectations and contributor behavior.
  • Templated review feedback that names a rejection as an approach difference should reduce the 37.02% of reported bias that stems from disagreements about style or design.
  • Since reviewer responsiveness is the top shared challenge, automation that pre-checks quality before human review should shorten the wait that most frustrates contributors.
  • Familiarity bias is the dominant real bias reported, so anonymized or blind review is the concrete intervention most likely to help newcomers and underrepresented contributors.
  • Developers who consult project documentation consistently rate the process as clearer and fairer, making documentation a low-cost lever for perceived bias.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the signal/noise split holds, future self-report studies should stop treating 'I experienced bias' as a single category and measure perceived unfairness separately from statistical discrimination.
  • The finding that maintainers report language challenges far more often than contributors suggests some 'approach differences' may really be fluency asymmetries; a testable extension is comparing review turnaround and revision counts for non-native-English contributors in the same repositories.
  • The fact that respondents were experienced, mostly male developers implies the 65% familiarity-bias share may understate newcomers' exposure; trace data on first-time contributors' review times could test this without new surveys.
  • The paper's role-priority gaps point to a diagnostic for future work: measure contributor-maintainer goal alignment before designing onboarding or review tooling, rather than assuming bias is the main failure mode.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. This mixed-methods study surveys 289 OSS developers (102 maintainers, 187 contributors) from 81 GitHub repositories, interviews 23 of them, and analyzes perceptions of code review objectives, challenges, bias, and documentation. The paper reports that maintainers and contributors largely agree on review objectives, but differ in emphasis: maintainers emphasize alignment with project goals and rationale, while contributors overvalue novelty. It also claims that many self-reported bias experiences are actually 'approach differences' rather than genuine bias, with 37.02% of bias reports classified this way, and that familiarity bias disproportionately affects underrepresented and newer contributors. The authors position their contribution as the first study of alignment in OSS code review perceptions, and they provide a data supplement for the survey instruments and coding.

Significance. If the reported role-based differences are real, the paper provides useful evidence about where friction in OSS code review originates, and the recommendation to improve documentation and communication is actionable. The study's strengths include a comparatively large survey sample, a mixed-methods design with follow-up interviews, explicit pilot validation, reproduction of interview transcription by two tools, and a public data supplement. The headline comparisons (e.g., project-goal emphasis 31.37% vs. 13.37%, novelty 11.23% vs. 1.96%) are internally plausible and consistent with prior qualitative work. However, the central RQ2 claim of 'misinterpretation of approach differences as bias' rests on a definitional choice that is not acknowledged in the paper, and the quantitative reporting lacks effect sizes, confidence intervals, and a clear denominator for the bias-percentage figures. These issues affect the load-bearing parts of the paper but are addressable in revision.

major comments (5)
  1. [Section 4.2 and Introduction] The paper switches definitions of bias mid-argument. The introduction adopts the Tversky and Kahneman definition of bias as a systematic deviation from objective standards, norms, or rationality [80], but Section 4.2 classifies 'approach differences' as not bias using a narrower definition from Ford et al. [36] that requires unfair favoring or disfavoring based on characteristics such as gender, race, or perceived experience. Under the paper's own opening definition, a reviewer who consistently favors their own coding style and demands rewrites to match it can qualify as a systematic deviation, i.e., as bias. The examples quoted in Section 5.4 (C109, C6) describe exactly this pattern. The claim that 37.02% of bias reports are 'misunderstandings' is therefore not demonstrated; it is an artifact of narrowing the definition after the fact. Please either apply one definition consistently throughout, or analyze the data under both definitions and report the sensitivity of the 37.02% estimate to the definitional choice.
  2. [Section 3.5] The paper explicitly states that formal inter-coder reliability measures were not used because coding disagreements were resolved through discussion. That is a major limitation for RQ2 because the entire 'approach difference vs. bias' boundary is a single subjective coding judgment that supports the headline 37.02% estimate. Without reliability metrics (e.g., Cohen's kappa on a subset of responses), readers cannot distinguish robust categorization from idiosyncratic interpretation. The authors should either report reliability on a coded subsample or substantially temper the claim that approach differences are 'misunderstood' as bias.
  3. [Section 4.1, Table 4] The Spearman correlations in Table 4 appear to be computed on aggregate group percentages rather than on individual responses, and no sample size, confidence interval, or effect size is reported. The chi-square tests are reported only as p<0.05, so the reader cannot assess the magnitude of the 'subtle but significant' differences. Given that the paper's contribution is about the degree of alignment, please report correlation coefficients with 95% confidence intervals, effect sizes for the chi-square comparisons (e.g., Cramér's V), and the raw counts underlying the percentages. Without these, the size and precision of the alleged differences cannot be evaluated.
  4. [Section 5.1 vs. Table 5] There is a direct numeric contradiction in a central finding. Section 5.1 states that '15% of Maintainers identified alignment with project goals as a primary objective, only 6.49% of Contributors shared this view,' while Table 5 reports 31.37% for Maintainers and 13.37% for Contributors. These cannot both be correct. The reported numbers must be reconciled and verified, since the project-goal difference is one of the paper's key findings.
  5. [Section 4.2, Table 10] The percentages in Table 10 sum to more than 100% (65.40 + 37.02 + 16.61 + 8.65 + 7.27 = 134.95), and the text reports different subgroup figures for language challenges (13.10% Maintainers vs. 2.77% Contributors in the text, but 24.10% vs. 2.70% in the Key Finding). The paper does not specify the denominator for these percentages or state whether participants could report multiple bias categories. Please clarify the coding scheme, the denominator, and the non-exclusivity of categories, and correct the inconsistent subgroup numbers, because the 37.02% misattribution estimate depends on this precision.
minor comments (4)
  1. [Throughout] There are several typos and stylistic inconsistencies, e.g., 'inline with pervius studies' (Section 4.2), 'What's more' capitalized mid-sentence (Sections 3.1 and 6.1), and 'or' in Table 9's group label. A careful proofreading pass is needed.
  2. [Section 4.1, Table 4] The 'Responses with Deviation' column lists categories that contributed to the chi-square deviation, but the table does not indicate the direction of each deviation (e.g., whether the category was over- or under-represented in each group). Adding a sign or arrow would improve interpretability.
  3. [Section 3.3] The recruitment description says repositories were filtered to those with 'at least three Maintainers' and non-English projects were excluded, but the possible bias from this filter is acknowledged only briefly. Since the sample consists of very large, popular repositories (median stars 35,507), the generalizability claims in Section 6.1 could be strengthened by explicitly discussing how this sampling frame might affect the bias-prevalence estimates.
  4. [Section 4.2] The paragraph on language challenges switches between percentages that appear to be based on different subpopulations (e.g., 16.61% overall, 13.10% Maintainers, 2.77% Contributors, then 24.10% vs. 2.70% in the Key Finding). Even if these come from different questions, the text should state clearly which question/denominator each figure refers to.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: results are descriptive survey/interview statistics; the approach-difference classification is a disclosed definitional choice, not a fitted input or self-citation chain.

full rationale

The paper contains no fitted parameters, no equations whose outputs equal inputs, and no load-bearing self-citation chain. RQ1–RQ3 findings are descriptive percentages, chi-square tests, and Likert summaries computed directly from survey responses, interviews, and repository metadata; there is no step where a quantity is fit to a subset and then 'predicted' for a closely related quantity. The only arguable definitional issue is in Section 4.2/5.4, where responses describing technical approach disagreements are coded as 'Approach Difference' and then declared not to meet the Ford et al. [36] definition of bias; the 37.02% 'misunderstanding' estimate therefore depends on the authors' coding and on that normative definition, and Section 3.5 explicitly discloses that formal inter-coder reliability was not computed. That is a validity and interpretation concern, not circularity: the category was created from participants' own descriptions, and the non-bias judgment is an external normative application, not a term defined so as to make the conclusion true by construction. The self-citation [69] is a data-availability link, and [68] is a related-work citation; neither carries the argument. Hence no significant circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claims rest on self-reported perceptions and author-conducted open coding; no free parameters are fitted and no invented entities are introduced. The main assumptions are the validity of survey self-reports, the adequacy of the Chi-square grouping, and the appropriateness of the authors' thematic classification.

assumptions (5)
  • domain assumption Survey and interview self-reports are treated as accurate accounts of code review experiences.
    All RQ1, RQ2, and RQ3 measurements rely on participant memory and interpretation; no behavioral trace data are used to validate them. Introduced in Sections 3.1 and 3.5.
  • domain assumption Open coding performed by the first two authors, with disagreements resolved by discussion, yields valid categories without formal inter-coder reliability.
    Section 3.5 states this explicitly; the 37.02% misattribution figure and all theme percentages inherit this assumption.
  • domain assumption Respondents can be cleanly classified as Contributors or Maintainers by primary role despite role duality across projects.
    Section 3.3 recruitment and Section 6.1 acknowledge that maintainers also contribute elsewhere, but analyses treat roles as distinct.
  • standard math Statistical test assumptions (Spearman monotonicity, Chi-square expected frequencies at least 5) are met after grouping rare responses.
    Section 3.5 states these assumptions and the grouping procedure; the manuscript does not show the grouped contingency tables.
  • domain assumption The definition of bias taken from Ford et al. [36] is the correct standard for classifying approach differences as non-bias.
    Section 4.2 and Section 5.4 use this definition to declare 37.02% of reports as misattribution.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Are We on the Same Page? Examining Developer Perception Alignment in Open Source Code Reviews." pith.science (2026). https://pith.science/paper/LWKFDZSL

@misc{pith2026250418407,
  author       = {Pith},
  title        = {Pith review of: Are We on the Same Page? Examining Developer Perception Alignment in Open Source Code Reviews},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LWKFDZSL}},
  note         = {Machine review of arXiv:2504.18407}
}
read the original abstract

Code reviews are a critical aspect of open-source software (OSS) development, ensuring quality and fostering collaboration. This study examines perceptions, challenges, and biases in OSS code review processes, focusing on the perspectives of Contributors and Maintainers. Through surveys (n=289), interviews (n=23), and repository analysis (n=81), we identify key areas of alignment and disparity. While both groups share common objectives, differences emerge in priorities, e.g, with Maintainers emphasizing alignment with project goals while Contributors overestimated the value of novelty. Bias, particularly familiarity bias, disproportionately affects underrepresented groups, discouraging participation and limiting community growth. Misinterpretation of approach differences as bias further complicates reviews. Our findings underscore the need for improved documentation, better tools, and automated solutions to address delays and enhance inclusivity. This work provides actionable strategies to promote fairness and sustain the long-term innovation of OSS ecosystems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What Motivates Whom? A Survey of Newcomers to OSS and Experienced OSS Practitioners

    cs.SE 2026-07 conditional novelty 5.0 of 10

    Demographics and motivations correlate with OSS project-selection preferences, with distinct patterns for newcomers versus experienced practitioners in a 208-person survey.

Reference graph

Works this paper leans on

92 extracted references · 78 canonical work pages · cited by 1 Pith paper

  1. [80]

    Tversky and D

    A. Tversky and D. Kahneman. Judgment Under Uncertainty: Heuristics and Biases. Science, 185(4157), pp. 1124–1131, 1974

  2. [36]

    Ford, D., Smith, J., Guo, P., & Zimmermann, T. (2019). Be yond the Code: Prioritiz- ing Inclusivity in Software Development. In Proceedings of the 2019 IEEE/ACM 41st International Conference on Software Engineering (pp . 1107–1118)

  3. [1]

    Bianca Trinkenreich, Igor Wiese, Anita Sarma, Marco Ger osa, and Igor Stein- macher. 2022. Women’s Participation in Open Source Software: A Survey of the Literature. ACM Trans. Softw. Eng. Methodol. 31, 4, Article 81 (October 2022), 37 pages

  4. [2]

    On older adults in free/open source software: reflections of contrib utors and community leaders,

    J. L. Davidson, R. Naik, U. A. Mannan, A. Azarbakht and C. J ensen, "On older adults in free/open source software: reflections of contrib utors and community leaders, " 2014 IEEE Symposium on Visual Languages and Human -Centric Com- puting (VL/HCC), Melbourne, VIC, Australia, 2014, pp. 93-1 00

  5. [3]

    Veteran develop- ers’ contributions and motivations: An open source perspective,

    P. Morrison, R. Pandita, E. Murphy-Hill and A. McLaughli n, "Veteran develop- ers’ contributions and motivations: An open source perspective, " 2016 IEEE Sym- posium on Visual Languages and Human-Centric Computing (VL /HCC), Cam- bridge, UK, 2016, pp. 171-179

  6. [4]

    Expectations , outcomes, and challenges of modern code review

    Bacchelli, Alberto, and Christian Bird. “Expectations , outcomes, and challenges of modern code review. ” In Proceedings of the 2013 International Conference on Software Engineering, pp. 712-721. 2013

  7. [5]

    McDowell, I., & Newell, C. (2009). Principles of medical statistics. Oxford Uni- versity Press

  8. [6]

    Altman, D. G. (1990). Practical statistics for medical r esearch. CRC press

Show all 92 references
  1. [7]

    B. G. Glaser and A. L. Strauss. The Discovery of Grounded Theory: Strategies for Qualitative Research. Aldine de Gruyter, 1967

  2. [8]

    Benington, H. D. (1983). Production of large computer pr ograms. Proceedings of the ONR Symposium on Advanced Programming Methods for Digit al Comput- ers, pp. 15–27

  3. [9]

    H. Patel. 2023. An Insight on (SDLC) Software Devel- opment Lifecycle Process Models. Advance. Available at: https://advance.sagepub.com/users/511828/articles/705050-an-insight-on-sdlc-software-development-lifec ycle-process-models

  4. [10]

    Fatima, N., Nazir, S., & Chuprat, S. (2020). Knowledge S haring Factors for Mod- ern Code Review to Minimize Software Engineering Waste. Int ernational Jour- nal of Advanced Computer Science and Applications, 11(1)

  5. [11]

    W., Kula, R

    Chouchen, M., Ouni, A., Mkaouer, M. W., Kula, R. G., & Ino ue, K. (2021). What makes a code review useful to OpenDev developers? An empiric al investigation. Empirical Software Engineering, 26(3), 1–28

  6. [12]

    M., Yin, G., & Wang, T

    Zhang, Y., Wang, H. M., Yin, G., & Wang, T. (2019). Unders tanding the time to first response in GitHub pull requests. arXiv preprint arXiv :2304.08426. Avail- able at: https://ar5iv.org/

  7. [13]

    Huang, Y., Leach, K., Sharafi, Z., McKay, N., Santander, T., & Weimer, W. (2020). Biases and differences in code review using medical imaging a nd eye-tracking: genders, humans, and machines. Proceedings of the 28th ACM J oint Meeting on European Software Engineering Conference ...

  8. [14]

    A., Ernst, N

    Storey, M. A., Ernst, N. A., Williams, C., & Kalliamvako u, E. (2020). The who, what, how of software engineering research: a socio-techni cal framework. Em- pirical Software Engineering, 25, 4097-4129

  9. [15]

    Vasilescu, B., Posnett, D., Ray, B., van den Brand, M. G. J., Serebrenik, A., Filkov, V., & Devanbu, P. (2015). Gender and Tenure Diversity in GitH ub Teams. In Pro- ceedings of the 33rd Annual ACM Conference on Human Factors i n Computing Systems (pp. 3789–3798)

  10. [16]

    Terrell, J., Kofink, A., Middleton, J., Rainear, C., Mur phy-Hill, E., Parnin, C., & Stallings, J. (2016). Gender Differences and Bias in Open Sou rce: Pull Request Acceptance of Women Versus Men. PeerJ Computer Science, 3, e 111

  11. [17]

    C., & Bird, C

    Rigby, P. C., & Bird, C. (2013). Convergent Software Pee r Review Practices. In Proceedings of the 2013 ACM SIGSOFT International Symposiu m on Founda- tions of Software Engineering (pp. 202–212)

  12. [18]

    and Redmi les, D.F., 2015

    Steinmacher, I., Silva, M.A.G., Gerosa, M.A. and Redmi les, D.F., 2015. A system- atic literature review on the barriers faced by newcomers to open source soft- ware projects. Information and Software Technology, 59, pp .67-85

  13. [19]

    Visibility of Women in the Software De- velopment Life Cycle

    Seguel, Marlene Valeria Negrier, et al. "Visibility of Women in the Software De- velopment Life Cycle. " Journal of Computer Science and Technology 23.2 (2023): e15-e15

  14. [20]

    HaTe Detector: A Tool for Detecting and Correcting Harmful Terminology in Comput ing Artifacts,

    H. Winchester, E. Al Haque, A. Boyd and B. Johnson, "HaTe Detector: A Tool for Detecting and Correcting Harmful Terminology in Comput ing Artifacts, " 2023 IEEE Symposium on Visual Languages and Human-Centric C omputing (VL/HCC), Washington, DC, USA, 2023, pp. 245-248,

  15. [21]

    Shyamal Mishra and Preetha Chatterjee. 2024. Explorin g ChatGPT for Toxicity Detection in GitHub. In Proceedings of the 2024 ACM/IEEE 44t h International Conference on Software Engineering: New Ideas and Emerging Results (ICSE- NIER’24). Association for Computing Machinery, Ne...

  16. [22]

    Singh, V., Brandon, W. (2019). Open Source Software Com munity Inclusion Ini- tiatives to Support Women Participation. In: Bordeleau, F. , Sillitti, A., Meirelles, P., Lenarduzzi, V. (eds) Open Source Systems. OSS 2019. IFIP Advances in Infor- mation and Communication Technolo...

  17. [23]

    and Storey, M.A., 2021

    Albusays, K., Bjorn, P., Dabbish, L., Ford, D., Murphy- Hill, E., Serebrenik, A. and Storey, M.A., 2021. The diversity crisis in software develo pment. IEEE Software, 38(2), pp.19-25

  18. [24]

    Nafus, D. (2012). ‘Patches don’t have gender’: What is n ot open in open source software. New Media & Society, 14(4), 669-683

  19. [25]

    Tsay, J., Dabbish, L., & Herbsleb, J. (2014). Influence o f Social and Technical Fac- tors for Evaluating Contribution in GitHub. In Proceedings of the 36th Interna- tional Conference on Software Engineering (pp. 356–366)

  20. [26]

    Egelman, Emerson Murphy-Hill, Elizabeth Ka mmer, Margaret Mor- row Hodges, Collin Green, Ciera Jaspan, and James Lin

    Carolyn D. Egelman, Emerson Murphy-Hill, Elizabeth Ka mmer, Margaret Mor- row Hodges, Collin Green, Ciera Jaspan, and James Lin. 2020. Predicting devel- opers’ negative feelings about code review. In Proceedings of the ACM/IEEE 42nd International Conference on Software Enginee...

  21. [27]

    Emerson Murphy-Hill, Ciera Jaspan, Carolyn Egelman, a nd Lan Cheng. 2022. The pushback effects of race, ethnicity, gender, and age in code review. Commun. ACM 65, 3 (March 2022), 52–57

  22. [28]

    S., & Bird, C

    Bosu, A., Greiler, M. S., & Bird, C. (2016). Characteris tics of Useful Code Reviews: An Empirical Study at Microsoft. In Proceedings of the 12th Working Conference on Mining Software Repositories (pp. 146–156)

  23. [29]

    Dabbish, L., Stuart, C., Tsay, J., & Herbsleb, J. (2012) . Social Coding in GitHub: Transparency and Collaboration in an Open Software Reposit ory. In Proceed- ings of the ACM 2012 Conference on Computer Supported Cooper ative Work (pp. 1277–1286)

  24. [30]

    Khalid, S., Brown, C. (2024). Exploring Stakeholder Ch allenges in Recruitment for Human-Centric Computing Research, 2024 IEEE Symposium on Visual Lan- guages and Human-Centric Computing (VL/HCC), Liverpool, U K, 2024

  25. [31]

    M., Adams, B., & Hassan, A

    Jiang, Z. M., Adams, B., & Hassan, A. E. (2013). Examinin g the Evolution of Code Review Processes in Mozilla Firefox. In 2013 IEEE Internati onal Conference on Software Maintenance (pp. 358–367)

  26. [32]

    Kononenko, O., Baysal, O., & Holmes, R. (2016). Mining M odern Code Review Repositories for Empirical Evidence. In Empirical Softwar e Engineering, 21(3), 1106–1152

  27. [33]

    McIntosh, S., Kamei, Y., Adams, B., & Hassan, A. E. (2016 ). An Empirical Study of the Impact of Modern Code Review Practices on Software Qua lity. Empirical Software Engineering, 21(5), 2146–2189

  28. [34]

    Diverse teams can improve engineering outco mes — but recent affirmative action decision may hinder efforts to create diver se teams

    Dowler, L. Diverse teams can improve engineering outco mes — but recent affirmative action decision may hinder efforts to create diver se teams. The Conversation. 2023. https://theconversation.com/diver se-teams-can-improve- engineering-outcomes-but-recent-affirmative-action-decisi...

  29. [35]

    and Marczak, S., 2018, October

    Toscani, C., Gery, D., Steinmacher, I. and Marczak, S., 2018, October. A gamifica- tion proposal to support the onboarding of newcomers in the fl osscoach portal. In Proceedings of the 17th Brazilian Symposium on Human Fact ors in Comput- ing Systems (pp. 1-10)

  30. [37]

    Bird, C., & Bacchelli, A. (2015). The Practices and Chal lenges of Modern Code Review. In IEEE Software, 32(6), 37–43

  31. [38]

    McGrath, S., & Zhou, W. (2019). Mitigating Biased Code R eviews in Open Source Projects. In Proceedings of the 2019 Conference on Fairness, Accountability, and Transparency (pp. 276–284)

  32. [39]

    O’Regan, J., Krishnamurthy, S., & Watson, R. T. (2021). A Study of Bias in Code Review Decisions in Open Source Projects. In Proceedings of the 2021 ACM SIGSOFT International Symposium on Foundations of Software Engineering (pp. 589–598)

  33. [40]

    Miku Watanabe, Yutaro Kashiwa, Bin Lin, Toshiki Hirao, Ken’Ichi Yamaguchi, and Hajimu Iida. 2024. On the Use of ChatGPT for Code Review: D o Developers Like Reviews By ChatGPT? In Proceedings of the 28th International Conference on Evaluation and Assessment in Software Enginee...

  34. [41]

    Lee, J., Xiao, Z., & Chan, W. K. (2020). Gender Participa tion Challenges in Open Source Software: A Multi-case Study. In Proceedings of the 2020 IEEE/ACM 42nd International Conference on Software Engineering (pp. 170 –181)

  35. [42]

    Eghbal, N. (2016). The Universe of Open Source Software : Where Did It All Start? In Open Source Software Report

  36. [43]

    Huang, L., Jackson, D., & Mahajan, D. (2018). Gender Dyn amics and Bias in Open Source Software Communities. In Empirical Software En gineering, 23(1), 128–152

  37. [44]

    Diversity and Inclusion Report

    GitHub (2021). Diversity and Inclusion Report. GitHub , Inc. Retrieved from https://github.blog/2021-09-08-diversity-inclusion-report-2021

  38. [45]

    Diversity and Inclusion Ini- tiatives in Open Source Communities

    Linux Foundation (2019). Diversity and Inclusion Ini- tiatives in Open Source Communities. Retrieved from https://www.linuxfoundation.org/diversity-initiatives

  39. [46]

    S., Anjana, A

    Adapa, C., Avulamanda, S. S., Anjana, A. R. K., & Victor, A. (2024). AI-Powered Code Review Assistant for Streamlining Pull Request Mergin g. In 2024 IEEE In- ternational Conference for Women in Innovation, Technolog y & Entrepreneur- ship (ICWITE), 323–327

  40. [47]

    Ahmed, T., Bosu, A., Iqbal, A., & Rahimi, S. (2017). Sent iCR: A Customized Sen- timent Analysis Tool for Code Review Interactions. In 2017 3 2nd IEEE/ACM International Conference on Automated Software Engineeri ng (ASE), 106–111. EASE 2025, 17–20 June, 2025, Istanbul, Türkiye...

  41. [48]

    Cohen, S. (2021). Contextualizing Toxicity in Open Sou rce: A Qualitative Study. In Proceedings of the 29th ACM Joint Meeting on European Soft ware Engineer- ing Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE), 1669–1671

  42. [49]

    Ferreira, I., Cheng, J., & Adams, B. (2021). The ’Shut th e F**k up’ Phenomenon: Characterizing Incivility in Open Source Code Review Discussions. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2), 353:1– 353:35

  43. [50]

    Sarker, A

    J. Sarker, A. K. Turzo, M. Dong, and A. Bosu. 2023. Automa ted Identification of Toxic Code Reviews Using ToxiCR. ACM Trans. Softw. Eng. Me thodol. 32, 5, Article 118 (2023). DOI: https://doi.org/10.1145/358356 2

  44. [51]

    Miller, C., Cohen, S., Klug, D., Vasilescu, B., & Kastne r, C. (2022). ’Did You Miss My Comment or What?’: Understanding Toxicity in Open Source Discussions. In Proceedings of the 44th International Conference on Soft ware Engineering

  45. [52]

    B. T. R. Savarimuthu, Z. Zareen, J. Cheriyan, M. Yasir, a nd M. Galster. 2023. Bar- riers for Social Inclusion in Online Software Engineering Communities: A Study of Offensive Language Use in Gitter Projects. In Proceedings of the 27th Interna- tional Conference on Evaluation a...

  46. [53]

    private- collective

    Hippel, E. V., & Krogh, G. V. (2003). Open source softwar e and the “private- collective” innovation model: Issues for organization sci ence. Organization sci- ence, 14(2), 209-223

  47. [54]

    Wh y Do Episodic Volunteers Stay in FLOSS Communities?,

    A. Barcomb, K. -J. Stol, D. Riehle and B. Fitzgerald, "Wh y Do Episodic Volunteers Stay in FLOSS Communities?, " 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE), Montreal, QC, Canada, 2019, p p. 948-959

  48. [55]

    Do a s I Do, Not as I Say: Do Contribution Guidelines Match the GitHub Contribution Process?,

    O. Elazhary, M. -A. Storey, N. Ernst and A. Zaidman, "Do a s I Do, Not as I Say: Do Contribution Guidelines Match the GitHub Contribution Process?, " 2019 IEEE In- ternational Conference on Software Maintenance and Evolution (ICSME), Cleve- land, OH, USA, 2019, pp. 286-290,

  49. [56]

    Barriers Faced by Women in Software Development Project s

    Dias Canedo E, Acco Tives H, Bogo Marioti M, Fagundes F, Siqueira de Cerqueira JA. Barriers Faced by Women in Software Development Project s. Information. 2019; 10(10):309

  50. [57]

    (2021, May)

    Dias, E., Meirelles, P., Castor, F., Steinmacher, I., Wiese, I., & Pinto, G. (2021, May). What makes a great maintainer of open source projects?. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE)

  51. [58]

    Justin Middleton, Emerson Murphy-Hill, Demetrius Gre en, Adam Meade, Roger Mayer, David White, and Steve McDonald. 2018. Which contrib utions predict whether developers are accepted into github teams. In Proce edings of the 15th International Conference on Mining Software Repo...

  52. [59]

    Setting guidelines for repository contributo rs

    GitHub. Setting guidelines for repository contributo rs. GitHub Docs. https://docs.github.com/en/communities/setting-up-your-project-for-healthy- contributions/setting-guidelines-for-repository-contributors

  53. [60]

    Shepherd, Igor Wiese, Chri stoph Treude, Marco Au- rélio Gerosa, and Igor Steinmacher

    Felipe Fronchetti, David C. Shepherd, Igor Wiese, Chri stoph Treude, Marco Au- rélio Gerosa, and Igor Steinmacher. 2023. Do CONTRIBUTING F iles Provide In- formation about OSS Newcomers’ Onboarding Barriers? In Pro ceedings of the 31st ACM Joint European Software Engineering C...

  54. [61]

    D., Weingart, L

    Murphy-Hill, E., Dicker, J., Horvath, A., Morrow Hodge s, M., Egelman, C. D., Weingart, L. R., Jaspan, C., Green, C., & Chen, N. (2023). Sys temic Gender In- equities in Who Reviews Code. Proceedings of the ACM on Human -Computer Interaction, 7(CSCW1), 1–59

  55. [62]

    Pavlína Wurzel Gonçalves, Gül Çalikli, and Alberto Bac chelli. 2022. Interper- sonal Conflicts During Code Review: Developers’ Experienceand Practices. Proc. ACM Hum.-Comput. Interact. 6, CSCW1, Article 98 (April 2022 ), 33 pages

  56. [63]

    Perception and misperception of bias i n human judgment

    Pronin, Emily. "Perception and misperception of bias i n human judgment. " Trends in cognitive sciences 11.1 (2007): 37-43

  57. [64]

    Evalu ating how static analysis tools can reduce code review effort,

    D. Singh, V. R. Sekar, K. T. Stolee and B. Johnson, "Evalu ating how static analysis tools can reduce code review effort, " 2017 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), Raleigh, NC, USA, 20 17, pp. 101-105

  58. [65]

    James Dominic, Jada Houser, Igor Steinmacher, Charles Ritter, and Paige Rodeghero. 2020. Conversational Bot for Newcomers Onboard ing to Open Source Projects. In Proceedings of the IEEE/ACM 42nd International Conference on Software Engineering Workshops (ICSEW’20). Associatio ...

  59. [66]

    “STILL AROUND

    S. Van Breukelen, A. Barcombt, S. Baltes and A. Serebren ik, "“STILL AROUND”: Experiences and Survival Strategies of Veteran Women Softw are Developers, " 2023 IEEE/ACM 45th International Conference on Software En gineering (ICSE), Melbourne, Australia, 2023, pp. 1148-1160

  60. [67]

    DevGPT: St udying Developer- ChatGPT Conversations,

    T. Xiao, C. Treude, H. Hata and K. Matsumoto, "DevGPT: St udying Developer- ChatGPT Conversations, " 2024 IEEE/ACM 21st InternationalConference on Min- ing Software Repositories (MSR), Lisbon, Portugal, 2024, p p. 227-230

  61. [68]

    an d Liebel, G., 2024

    Hyrynsalmi, S.M., Baltes, S., Brown, C., Prikladnicki , R., Rodriguez-Perez, G., Serebrenik, A., Simmonds, J., Trinkenreich, B., Wang, Y. an d Liebel, G., 2024. Bridging Gaps, Building Futures: Advancing Software Devel oper Diversity and Inclusion Through Future-Oriented Resea...

  62. [69]

    Are We on the Same Page? Examining Developer P erception Alignment in Open Source Code Reviews

    Alebachew, Yoseph Berhanu; Ko, Minhyuk; Brown, Dwayne (2025). Data as- sociated with "Are We on the Same Page? Examining Developer P erception Alignment in Open Source Code Reviews". University Librari es, Virginia Tech. Dataset. https://doi.org/10.7294/28816835

  63. [70]

    Open Source Software and the ’Private- Collective’ Innovation Model: Issues for Organization Science,

    E. von Hippel and G. von Krogh, “Open Source Software and the ’Private- Collective’ Innovation Model: Issues for Organization Science, ”Organization Sci- ence, vol. 14, no. 2, pp. 209–223, 2003

  64. [71]

    A Framework for Understan ding Open Source Software Development,

    J. Feller and B. Fitzgerald, “A Framework for Understan ding Open Source Software Development, ” in Proceedings of the 21st International Con- ference on Information Systems (ICIS 2000) , 2000, pp. 58–69. Available: https://aisel.aisnet.org/icis2000/58/

  65. [72]

    The proof and measurement of association between two things,

    C. Spearman, “The proof and measurement of association between two things, ” The American Journal of Psychology , vol. 15, no. 1, pp. 72–101, 1904

  66. [73]

    J. D. Gibbons and S. Chakraborti, Nonparametric Statistical Inference, 3rd ed. New York: Marcel Dekker, 1993

  67. [74]

    K. Pearson, “On the criterion that a given system of devi ations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling, ” The London, Edinburgh, and Dublin Philosophical Magazine and Jou...

  68. [75]

    Agresti, Statistical Methods for the Social Sciences , 5th ed

    A. Agresti, Statistical Methods for the Social Sciences , 5th ed. New York: Pearson, 2018

  69. [76]

    The analysis of residuals in cross-cla ssified tables,

    S. J. Haberman, “The analysis of residuals in cross-cla ssified tables, ”Biometrics, vol. 29, no. 1, pp. 205–220, 1973

  70. [77]

    Sultana, A

    S. Sultana, A. K. Turzo, and A. Bosu, “Code reviews in ope n source projects: how do gender biases affect participation and outcomes?, Empirical Software En- gineering, vol. 28, no. 92, pp. 1–30, 2023

  71. [78]

    A. G. Greenwald and M. R. Banaji. Implicit Social Cognition: Attitudes, Self- Esteem, and Stereotypes. Psychological Review, 102(1), pp. 4–27, 1995

  72. [79]

    Lerner and J

    J. Lerner and J. Tirole. Some Simple Economics of Open Source . The Journal of Industrial Economics, 50(2), pp. 197–234, 2002

  73. [81]

    W., & Nunes, I

    Santos, E. W., & Nunes, I. (2022). Investigating the effe ctiveness of peer code review in distributed software development based on object ive and subjective data. Empirical Software Engineering, 27(13)

  74. [82]

    U., & Carver, J

    Eisty, N. U., & Carver, J. C. (2022). Developers’ percep tion of peer code review in research software development. Empirical Software Engi neering, 27(13)

  75. [83]

    Dias, E., Meirelles, P., Castor, F., Steinmacher, I., W iese, I., & Pinto, G. (2021). What Makes a Great Maintainer of Open Source Projects? In IEE E/ACM Inter- national Conference on Software Engineering (ICSE), 982–9 94

  76. [84]

    Constantino, K., Souza, M., Zhou, S., Figueiredo, E., & Kästner, C. (2021). Percep- tions of open-source software developers on collaboration s: An interview and survey study. Empirical Software Engineering, 26, 1–33

  77. [85]

    Wessel, M., Serebrenik, A., & Wiese, I. (2021). What to E xpect from Code Review Bots on GitHub? A Survey with OSS Maintainers. In ACM/IEEE In ternational Conference on Software Engineering (ICSE), 1134–1145

  78. [86]

    Elazhary, O., Storey, M.-A., Ernst, N., & Zaidman, A. (2 022). Do as I Do, Not as I Say: Do Contribution Guidelines Match the GitHub Contribu tion Process? In IEEE/ACM International Conference on Software Engineerin g (ICSE), 942–954

  79. [87]

    Beller, M., Bacchelli, A., Zaidman, A., & Juergens, E. ( 2021). Modern code re- views in open-source projects: Which problems do they fix? Em pirical Software Engineering, 26(1), 1–24

  80. [88]

    L., & Wasowski, A

    Alami, A., Cohn, M. L., & Wasowski, A. (2020). Why Does Co de Review Work for Open Source Software Communities? Empirical Software Engineering, 25(1), 1–18

  81. [89]

    B., & Hanssen, G

    Stray, V., Moe, N. B., & Hanssen, G. K. (2020). Coordinat ion in Global Software Development: Challenges and Coordination Mechanisms. Jou rnal of Systems and Software, 168, 110673

  82. [90]

    Improving de- veloper participation rates in surveys,

    E. Smith, R. Loftin, E. Murphy-Hill, C. Bird, and T. Zimm ermann, "Improving de- veloper participation rates in surveys, " in *Proceedings o f the 2013 6th Interna- tional Workshop on Cooperative and Human Aspects of Softwar e Engineering (CHASE)*, 2013, pp. 89–92, IEEE

  83. [91]

    A systematic literature review on the b arriers faced by new- comers to open source software projects

    Igor Steinmacher, Marco Aurelio Graciotto Silva, Marc o Aurelio Gerosa , David F. Redmiles (2015). "A systematic literature review on the b arriers faced by new- comers to open source software projects" Information and Software Technology, Vol 59, 67-85

  84. [92]

    Avelino, G., Passos, L., Hora, A., & Valente, M. T. (2019 ). A novel approach for estimating truck factors. Empirical Software Engineering , 24(1), 1-33. Received 31 January 2025

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.