{"id":"5935a546-1bf3-407a-a548-6b92b76925d5","arxiv_id":"2504.18407","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Contributors and maintainers broadly agree on code review goals, but contributors overvalue novelty and undervalue rationale while maintainers emphasize project alignment.","lead":"This study surveyed 289 open source developers and interviewed 23 to compare how contributors and maintainers see code reviews. It found mostly shared goals but real differences in priorities, and that many perceived biases are actually disagreements about approach.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.2 switches definitions of bias mid-argument: the paper opens with Tversky–Kahneman 'systematic deviation' but labels 'approach difference' as non-bias using Ford et al.'s narrower group-based definition, so the 37.02% misattribution claim rests on an unacknowledged definitional choice.","rationale":"The reader identified the coding of self-reported descriptions and the normative definition from Ford et al. as the weakest assumption; I agree, and I want to sharpen why the definitional issue is the most load-bearing. The paper's headline contribution is 'empirical proof of misattributing approach differences as bias.' That claim depends entirely on classifying the 'approach difference' category as non-bias. But the paper gives two incompatible definitions of bias: one in the introduction (Tversky and Kahneman, broad) and one in Section 4.2 (Ford et al., narrow, group-based). Under the broad definition, a reviewer's consistent preference for a particular coding style or technique is a systematic deviation from objective standards and can be bias. Under the narrow definition, it is not. The authors use the narrow definition without acknowledging that this choice determines the result. This is an internal inconsistency, not merely a disagreement with an external consensus. The absence of inter-coder reliability compounds the issue: even under a fixed definition, the boundary between 'approach difference' and 'bias' is subjective, and the paper states that disagreements were resolved by discussion rather than measured. A re-coding check with independent raters would settle whether the category is stable and whether the 37.02% estimate survives application of the paper's own opening definition. I would keep the overall verdict conditional rather than rejecting the paper, because the survey data and role-based differences in Tables 5 and 6 may still be informative, but the 'misunderstanding of bias' framing and the abstract's 'empirical proof' language should be tempered until the definitional and coding concerns are addressed.","tokens_in":21849,"tokens_out":6990,"duration_ms":69481,"concrete_test":"Using the published supplementary data [69], have two independent raters blind to the authors' Table 10 labels and to the participant role apply the paper's own §1 definition of bias (Tversky and Kahneman [80], systematic deviation from objective standards/norms/rationality) to every excerpt coded as 'Approach Difference'; compute inter-rater agreement (Cohen's kappa) and the fraction meeting that definition. If the fraction exceeds 25% or kappa falls below 0.6, the 37.02% misattribution estimate and the RQ2 narrative need to be revised or heavily qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central RQ2 claim—that 37.02% of self-reported bias experiences are actually 'approach differences' and therefore not bias—depends on which definition of bias is applied. The introduction adopts the Tversky and Kahneman [80] definition: bias as 'a systematic deviation from objective standards, norms, or rationality in judgment or decision-making.' Section 4.2 then silently switches to Ford et al. [36]: unfair favoring or disfavoring based on characteristics unrelated to code quality, such as gender, race, or perceived experience. Under the paper's own opening definition, a reviewer who consistently favors their own coding style or demands that the author rewrite to the reviewer's preferred approach can qualify as a systematic deviation, i.e., as bias. The examples quoted in §5.4 (C109, C6) describe exactly this pattern. Therefore 'misunderstanding of bias' is not demonstrated; it is an artifact of narrowing the definition after the fact. The absence of inter-coder reliability, explicitly stated in §3.5, further weakens the category boundary, because 'approach difference' versus 'bias' is the one classification on which the headline estimate rests.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This mixed-methods study surveys 289 OSS developers (102 maintainers, 187 contributors) from 81 GitHub repositories, interviews 23 of them, and analyzes perceptions of code review objectives, challenges, bias, and documentation. The paper reports that maintainers and contributors largely agree on review objectives, but differ in emphasis: maintainers emphasize alignment with project goals and rationale, while contributors overvalue novelty. It also claims that many self-reported bias experiences are actually 'approach differences' rather than genuine bias, with 37.02% of bias reports classified this way, and that familiarity bias disproportionately affects underrepresented and newer contributors. The authors position their contribution as the first study of alignment in OSS code review perceptions, and they provide a data supplement for the survey instruments and coding.","tokens_in":22118,"tokens_out":3477,"duration_ms":32850,"significance":"If the reported role-based differences are real, the paper provides useful evidence about where friction in OSS code review originates, and the recommendation to improve documentation and communication is actionable. The study's strengths include a comparatively large survey sample, a mixed-methods design with follow-up interviews, explicit pilot validation, reproduction of interview transcription by two tools, and a public data supplement. The headline comparisons (e.g., project-goal emphasis 31.37% vs. 13.37%, novelty 11.23% vs. 1.96%) are internally plausible and consistent with prior qualitative work. However, the central RQ2 claim of 'misinterpretation of approach differences as bias' rests on a definitional choice that is not acknowledged in the paper, and the quantitative reporting lacks effect sizes, confidence intervals, and a clear denominator for the bias-percentage figures. These issues affect the load-bearing parts of the paper but are addressable in revision.","major_comments":[{"comment":"The paper switches definitions of bias mid-argument. The introduction adopts the Tversky and Kahneman definition of bias as a systematic deviation from objective standards, norms, or rationality [80], but Section 4.2 classifies 'approach differences' as not bias using a narrower definition from Ford et al. [36] that requires unfair favoring or disfavoring based on characteristics such as gender, race, or perceived experience. Under the paper's own opening definition, a reviewer who consistently favors their own coding style and demands rewrites to match it can qualify as a systematic deviation, i.e., as bias. The examples quoted in Section 5.4 (C109, C6) describe exactly this pattern. The claim that 37.02% of bias reports are 'misunderstandings' is therefore not demonstrated; it is an artifact of narrowing the definition after the fact. Please either apply one definition consistently throughout, or analyze the data under both definitions and report the sensitivity of the 37.02% estimate to the definitional choice.","section":"Section 4.2 and Introduction"},{"comment":"The paper explicitly states that formal inter-coder reliability measures were not used because coding disagreements were resolved through discussion. That is a major limitation for RQ2 because the entire 'approach difference vs. bias' boundary is a single subjective coding judgment that supports the headline 37.02% estimate. Without reliability metrics (e.g., Cohen's kappa on a subset of responses), readers cannot distinguish robust categorization from idiosyncratic interpretation. The authors should either report reliability on a coded subsample or substantially temper the claim that approach differences are 'misunderstood' as bias.","section":"Section 3.5"},{"comment":"The Spearman correlations in Table 4 appear to be computed on aggregate group percentages rather than on individual responses, and no sample size, confidence interval, or effect size is reported. The chi-square tests are reported only as p<0.05, so the reader cannot assess the magnitude of the 'subtle but significant' differences. Given that the paper's contribution is about the degree of alignment, please report correlation coefficients with 95% confidence intervals, effect sizes for the chi-square comparisons (e.g., Cramér's V), and the raw counts underlying the percentages. Without these, the size and precision of the alleged differences cannot be evaluated.","section":"Section 4.1, Table 4"},{"comment":"There is a direct numeric contradiction in a central finding. Section 5.1 states that '15% of Maintainers identified alignment with project goals as a primary objective, only 6.49% of Contributors shared this view,' while Table 5 reports 31.37% for Maintainers and 13.37% for Contributors. These cannot both be correct. The reported numbers must be reconciled and verified, since the project-goal difference is one of the paper's key findings.","section":"Section 5.1 vs. Table 5"},{"comment":"The percentages in Table 10 sum to more than 100% (65.40 + 37.02 + 16.61 + 8.65 + 7.27 = 134.95), and the text reports different subgroup figures for language challenges (13.10% Maintainers vs. 2.77% Contributors in the text, but 24.10% vs. 2.70% in the Key Finding). The paper does not specify the denominator for these percentages or state whether participants could report multiple bias categories. Please clarify the coding scheme, the denominator, and the non-exclusivity of categories, and correct the inconsistent subgroup numbers, because the 37.02% misattribution estimate depends on this precision.","section":"Section 4.2, Table 10"}],"minor_comments":[{"comment":"There are several typos and stylistic inconsistencies, e.g., 'inline with pervius studies' (Section 4.2), 'What's more' capitalized mid-sentence (Sections 3.1 and 6.1), and 'or' in Table 9's group label. A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The 'Responses with Deviation' column lists categories that contributed to the chi-square deviation, but the table does not indicate the direction of each deviation (e.g., whether the category was over- or under-represented in each group). Adding a sign or arrow would improve interpretability.","section":"Section 4.1, Table 4"},{"comment":"The recruitment description says repositories were filtered to those with 'at least three Maintainers' and non-English projects were excluded, but the possible bias from this filter is acknowledged only briefly. Since the sample consists of very large, popular repositories (median stars 35,507), the generalizability claims in Section 6.1 could be strengthened by explicitly discussing how this sampling frame might affect the bias-prevalence estimates.","section":"Section 3.3"},{"comment":"The paragraph on language challenges switches between percentages that appear to be based on different subpopulations (e.g., 16.61% overall, 13.10% Maintainers, 2.77% Contributors, then 24.10% vs. 2.70% in the Key Finding). Even if these come from different questions, the text should state clearly which question/denominator each figure refers to.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant topic for EASE and brings a sizable mixed-methods dataset. My main concern is that the RQ2 interpretation—the 'misunderstanding of bias' claim—is presented as an empirical result even though it depends on an unacknowledged definitional choice and on a single coding decision without reliability evidence. The numeric inconsistencies in Section 5.1 and Section 4.2 should also be caught before any acceptance decision. None of these issues appear to require new data collection; they can be fixed with a careful reanalysis and transparent reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The paper delivers the first direct survey comparison of how OSS contributors and maintainers perceive code review objectives, and the headline result is credible: maintainers emphasize project-goal alignment (31.37% vs 13.37%), contributors overvalue novelty (11.23% vs 1.96%) and underrate the need for rationale (3.74% vs 12.75%). That's a useful, actionable finding. The second thing is that the paper's other big claim—that 37.02% of perceived bias is actually just 'approach difference' and therefore not bias—is shakier than the abstract suggests, because it relies on swapping definitions of bias mid-argument.\n\nWhere credit is due: the study is well-designed for its purpose. The mixed-methods setup (n=289 survey, 23 interviews) is solid; the sample is large for OSS studies; they provide a dataset via a Virginia Tech DOI; and the qualitative quotes illustrate the quantitative gaps well. The RQ1/RQ3 material on documentation is a nice addition.\n\nSoft spots, in proportion. The big one is the definitional switch: the intro adopts Tversky–Kahneman's broad 'systematic deviation' definition, but Section 4.2 reclassifies 'approach differences' as non-bias using Ford et al.'s narrow group-based definition. Under the paper's own opening definition, a reviewer who systematically favors their own coding style can be biased. The examples quoted in Section 5.4 (C109, C6) read exactly that way. So the 'misunderstanding of bias' conclusion isn't demonstrated—it's an artifact of narrowing the definition after the fact. The paper needs to either commit to one definition or explicitly discuss how the results change under each. Related: the coding of the categories that drive that 37.02% estimate was done without inter-coder reliability, and the paper hand-waves it in Section 3.5. That's a minor issue for RQ1 themes but load-bearing for RQ2. Also, the abstract overclaims: there's no real repository analysis despite the abstract saying so—the 'n=81' is just the projects the participants come from—and the claim that familiarity bias 'disproportionately affects underrepresented groups' isn't supported by the data presented (they couldn't analyze race/ethnicity at all, and the significant difference was by role, not underrepresented status). Finally, the stats are bare-bones: chi-square with no effect sizes or confidence intervals, and the Spearman correlations are computed on aggregate percentages.\n\nOverall: the empirical core holds up for the alignment findings, and the bias-misattribution claim needs rework before it's taken at face value. The paper is worth a serious referee's time—I'd send it out and ask for major revision focused on definitional clarity, effect sizes, and aligning the abstract with the evidence. If you work on OSS or bias in collaboration, it's worth a read.","headline":"First survey comparison of contributor/maintainer review perceptions has a credible core, but the 'misperceived bias' claim rests on a definitional switch the authors don't acknowledge.","tokens_in":22624,"tokens_out":2995,"would_cite":true,"duration_ms":28227,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Open-source code-review friction comes less from bias than from contributors and maintainers valuing different things—maintainers want project fit, contributors lead with novelty.","keywords":["code review","open source software","perception alignment","bias","contributors and maintainers","familiarity bias","contribution guidelines","mixed methods"],"falsifier":"Have two independent teams, blind to the paper's categories, re-code the same open-ended bias descriptions; if coders cannot agree, or if the same incident is placed in both 'familiarity bias' and 'approach difference,' the 37.02% misattribution figure has no stable referent. A second check: in follow-up interviews, ask participants whether they would still call the incident unfair once the reviewer's preference is explained as a technical style preference—if they do, the paper's signal/noise split does not capture how bias is experienced.","tokens_in":21673,"feed_emoji":"⚖️","tokens_out":9955,"duration_ms":92423,"temperature":0.7,"pith_summary":"This paper argues that contributors and maintainers in open-source code review largely share one picture of what reviews should do, but systematically differ in emphasis, and that a large share of what developers call bias is really those emphasis gaps rather than discrimination. The evidence is a survey of 289 developers from 81 GitHub projects, interviews with 23 of them, and repository metadata; responses to identical questions are compared by role. Maintainers stress alignment with project goals, while contributors overvalue novelty and under-explain rationale. Among participants who reported bias, 65.40% described favoritism toward familiar contributors, while 37.02% described what the authors classify as differences in technical approach—experiences that do not meet the paper's working definition of bias. The paper draws practical consequences: better contribution guidelines, pre-review automation, and anonymized review should reduce both real bias and miscommunication.","feed_headline":"Many 'bias' reports in open source reviews are really goal mismatches","feed_subtitle":"Maintainers prize project fit; contributors prize novelty; familiarity bias is the real exception.","key_machinery":"The mechanism is a paired perception-alignment study: separate role-specific surveys ask contributors and maintainers the same open-ended questions about review objectives, acceptance factors, challenges, and improvements; open coding turns free text into theme percentages; Spearman correlation measures overall alignment; and chi-square post-hoc tests locate which themes deviate significantly. The same machinery is applied to bias: respondents who said they witnessed bias were asked to describe it, and the authors classified descriptions into five categories (familiarity bias, approach difference, language challenges, misunderstanding, other). The categories do the paper's work—they separate signal (true bias) from noise (perceived bias caused by mismatched expectations), and the tabulated percentages become the paper's evidence that a large portion of reported bias is misattribution rather than discrimination.","core_discovery":"The central claim is that perceived bias in open-source code review is a mix of signal and noise. The signal is real favoritism toward familiar contributors—named familiarity bias and reported by 65.40% of participants who noticed bias—which disproportionately hurts newer and underrepresented contributors. The noise consists of experiences the authors code as approach differences (37.02%), where a reviewer pushes for their preferred coding style, design, or technical solution; these are not unfair treatment under the definition the authors adopt, yet participants experience and describe them as bias. On the goals of review, maintainers and contributors agree on correctness, quality, standards, and documentation but diverge in emphasis: 31.37% of maintainers list alignment with project goals as an objective versus 13.37% of contributors, while 11.23% of contributors list novelty as a factor in pull-request acceptance versus 1.96% of maintainers, and only 3.74% of contributors mention rationale versus 12.75% of maintainers. The paper presents these role-based gaps in expectations as a cause of friction and disengagement that can be misread as bias.","pith_inferences":["If the signal/noise split holds, future self-report studies should stop treating 'I experienced bias' as a single category and measure perceived unfairness separately from statistical discrimination.","The finding that maintainers report language challenges far more often than contributors suggests some 'approach differences' may really be fluency asymmetries; a testable extension is comparing review turnaround and revision counts for non-native-English contributors in the same repositories.","The fact that respondents were experienced, mostly male developers implies the 65% familiarity-bias share may understate newcomers' exposure; trace data on first-time contributors' review times could test this without new surveys.","The paper's role-priority gaps point to a diagnostic for future work: measure contributor-maintainer goal alignment before designing onboarding or review tooling, rather than assuming bias is the main failure mode."],"forward_implications":["Writing project goals and the need for rationale directly into contribution guidelines should close the largest measurable gap between maintainer expectations and contributor behavior.","Templated review feedback that names a rejection as an approach difference should reduce the 37.02% of reported bias that stems from disagreements about style or design.","Since reviewer responsiveness is the top shared challenge, automation that pre-checks quality before human review should shorten the wait that most frustrates contributors.","Familiarity bias is the dominant real bias reported, so anonymized or blind review is the concrete intervention most likely to help newcomers and underrepresented contributors.","Developers who consult project documentation consistently rate the process as clearer and fairer, making documentation a low-cost lever for perceived bias."],"supporting_citations":[{"why":"Supplies the prior enumeration of modern code-review expectations and outcomes that the survey's objective and key-factor themes are compared with.","marker":"[4]"},{"why":"Provides the prior result that misunderstandings in GitHub contributions are often rooted in differences in expectations rather than actual bias, which this study extends to perceived bias.","marker":"[25]"},{"why":"Gives the formal definition of code-review bias used to classify approach differences as misattribution rather than unfair treatment.","marker":"[36]"},{"why":"Documents biased review outcomes against underrepresented groups, which the familiarity-bias finding is said to align with.","marker":"[38]"},{"why":"Adds supporting evidence that bias and participation barriers affect underrepresented contributors in open-source software.","marker":"[41]"},{"why":"Shows that contribution guidelines often fail to match actual GitHub contribution practices, the gap investigated in RQ3.","marker":"[55]"},{"why":"Supports the claim that people overestimate the frequency and impact of bias in their own judgments.","marker":"[63]"},{"why":"Supplies the foundational definition of bias as systematic deviation from norms or rationality that frames the study.","marker":"[80]"}],"fun_headline_variants":["Open source code review 'bias' often hides goal mismatches","Perceived bias in code reviews is mostly approach differences","Maintainers and contributors disagree on code review goals","Familiarity bias real, but much 'bias' is just goal clash","Code review friction: maintainers want fit, contributors want novelty"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the assumption that independent coders can consistently tell a real bias experience apart from a mere difference in coding style or design preference, and that participants' word 'bias' means the same thing the coders' formal definition assumes; the paper states that coding disagreements were resolved by discussion rather than measured, so this classification is the step most worth checking.","fun_headline_variants_meta":{"raw":{"variants":["Open source code review 'bias' often hides goal mismatches","Perceived bias in code reviews is mostly approach differences","Maintainers and contributors disagree on code review goals","Familiarity bias real, but much 'bias' is just goal clash","Code review friction: maintainers want fit, contributors want novelty"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000573,"raw_usage":{"total_tokens":2698,"prompt_tokens":929,"completion_tokens":1769,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":1682}},"tokens_in":545,"tokens_out":1769,"duration_ms":11147,"temperature":1.0,"reasoning_tokens":1682,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:16:24.160052+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent teams, blind to the paper's categories, re-code the same open-ended bias descriptions; if coders cannot agree, or if the same incident is placed in both 'familiarity bias' and 'approach difference,' the 37.02% misattribution figure has no stable referent. A second check: in follow-up interviews, ask participants whether they would still call the incident unfair once the reviewer's preference is explained as a technical style preference—if they do, the paper's signal/noise split does not capture how bias is experienced.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior result that misunderstandings in GitHub contributions are often rooted in differences in expectations rather than actual bias, which this study extends to perceived bias."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the formal definition of code-review bias used to classify approach differences as misattribution rather than unfair treatment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents biased review outcomes against underrepresented groups, which the familiarity-bias finding is said to align with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Adds supporting evidence that bias and participation barriers affect underrepresented contributors in open-source software."},{"cited_title":"Do a s I Do, Not as I Say: Do Contribution Guidelines Match the GitHub Contribution Process?,","cited_arxiv_id":null,"evidence_quote":"Shows that contribution guidelines often fail to match actual GitHub contribution practices, the gap investigated in RQ3."},{"cited_title":"Perception and misperception of bias i n human judgment","cited_arxiv_id":null,"evidence_quote":"Supports the claim that people overestimate the frequency and impact of bias in their own judgments."},{"cited_title":"Tversky and D","cited_arxiv_id":null,"evidence_quote":"Supplies the foundational definition of bias as systematic deviation from norms or rationality that frames the study."}],"review_version":1}