{"id":"73b3702a-5ffd-4a39-bcbd-547df29958e9","arxiv_id":"1908.01718","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Peer evaluators in a real firm show 'discriminatory generosity': they downgrade qualified same-rank coworkers and overrate unqualified ones, consistent with strategic manipulation in crowdsourced performance assessment.","lead":"Using five years of performance ratings from a Chinese professional service firm, this paper finds that employees give lower ratings to same-rank coworkers who have already met promotion requirements and higher ratings to same-rank coworkers who have not. The result matters because it quantifies strategic bias in crowdsourced peer assessment, a system used by many companies for pay and promotion decisions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The negative interaction coefficient cannot by itself identify strategic behavior; it is equally consistent with peers' private information about post-qualification performance or with nonpeer credential-anchoring.","rationale":"The reader's conditional verdict is appropriate. The negative interaction β3 is the entire basis for the headline 'discriminatory generosity' claim, and the paper's identification depends on assuming that peer, nonpeer, and head raters are comparable except for conflict of interest. The most serious unaddressed alternative is that peers are better informed: if qualified employees' current performance declines after passing objective requirements, or if nonpeers simply anchor on credentials, the same β3 arises with no strategic behavior. Eq. (1)'s fixed effects do not remove ratee-specific time-varying shocks or differential information. The paper's own statement that head ratings are 'a nonstrategic benchmark' (Section 4.2) is an assertion, not a test. However, the regression results are internally consistent, the descriptive patterns are clear, and the paper honestly labels the counterfactual simulations exploratory. The correct verdict is therefore conditional, not reject, because the key estimand is identified only under assumptions that can be tested with the authors' data. Re-estimating with head ratings as the baseline would directly probe the benchmark assumption; if the interaction survives, the strategic interpretation is much better supported, while if it does not, the central claim should be downgraded.","tokens_in":8173,"tokens_out":15408,"duration_ms":180431,"concrete_test":"Re-estimate Eq. (1) on the subsample consisting only of same-rank peer ratings and department-head ratings, with head ratings as the baseline (Peer=0). If β3 remains significant and close to -0.493, the result is robust to replacing the mixed nonpeer group with the paper's own 'nonstrategic benchmark' and the strategic interpretation is strengthened. If β3 attenuates toward zero, the headline claim is an artifact of comparing peers with uninformed or credential-anchored nonpeer raters, and the verdict should be lowered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference—strategic 'discriminatory generosity'—rests on β3 = -0.493 (SE 0.079) in Eq. (1), under the untested assumption that nonpeer and head raters are an unbiased benchmark and that Qual does not proxy for time-varying ratee performance. Section 4.2 asserts that 'Interpreting department heads' ratings as a nonstrategic benchmark' supports the story, but it does not establish that nonpeer raters observe current performance rather than rewarding formal qualifications. The fixed effects (ratee, ratee rank, department, year) absorb level differences and common shocks, not ratee-specific performance trends or differential information. Same-rank peers are precisely the raters most likely to observe a qualified employee's actual current performance; if that performance declines after objective requirements are passed, or if nonpeers merely anchor on Qual, the same negative interaction is predicted by honest rating. Since Qual is absorbing for most employees, within-ratee variation is a before/after comparison at the qualification event, and no leads/lags, placebo, or rater-pair analysis is presented. Thus β3≠0 does not by itself distinguish strategic manipulation from information asymmetry.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies bias and strategic behavior in crowdsourced performance assessment using a five-year HR archive from a Chinese professional service firm, with 7,778 rater-ratee-year observations. The main analysis estimates an OLS regression of ratings on a peer-rater indicator, an objective promotion-qualification indicator, and their interaction, together with ratee, ratee-rank, department, and year fixed effects (Eq. (1), Table 3). The key finding is a significantly negative interaction coefficient (β3 = -0.493, SE 0.079), which the authors interpret as 'discriminatory generosity': peer raters downgrade qualified peers and overrate unqualified peers, relative to nonpeer raters. The paper also documents self-assessment inflation (§4.1) and presents counterfactual simulations of promotion probabilities under alternative rating schemes (§5.2). The central claim is that the observed pattern constitutes strategic manipulation in peer evaluation.","tokens_in":8366,"tokens_out":3961,"duration_ms":40403,"significance":"If the causal interpretation holds, this is one of the first field-data demonstrations of strategic manipulation in crowdsourced performance reviews, and the interaction-based design using an objective qualification measure is a useful template for fairness-aware talent analytics. The paper's strengths include clean OLS estimation with detailed fixed effects, transparent reporting of coefficients and clustered standard errors, and the use of an external objective benchmark (Qual) that is estimated rather than imposed. The descriptive pattern is robust and interesting: the negative interaction is large, precisely estimated, and internally consistent with the self-assessment results and the counterfactual correlations. However, the central inference from β3 to strategic behavior rests on untested identifying assumptions; the causal label is not yet established by the evidence presented.","major_comments":[{"comment":"The interpretation of β3 = -0.493 as evidence of strategic manipulation requires that, conditional on ratee, rank, department, and year fixed effects, Qual is uncorrelated with ratee-specific time-varying performance and that nonpeer and department-head ratings are unbiased benchmarks. Neither assumption is tested. The negative interaction is equally consistent with peers using private information that qualified peers' current performance is lower after passing requirements, or with nonpeer raters anchoring on formal credentials without observing current performance. Section 4.2 asserts that 'interpreting department heads' ratings as a nonstrategic benchmark' supports the interpretation, but this is an assertion, not a test; the higher correlation between head and nonpeer ratings (0.827) than between head and peer ratings (0.520) does not establish that nonpeers are unbiased. The manuscript should provide additional evidence for the causal reading, such as leads/lags around the qualification event, placebo tests using future qualification status, rater-pair analyses, or a direct discussion of why the information-asymmetry interpretation is ruled out. Without such evidence, the headline claim of strategic 'discriminatory generosity' is overreach.","section":"§4.3, Eq. (1), Table 3"},{"comment":"The counterfactual simulation of promotion probabilities rests on the logit model in Table 4, which itself includes the endogenous PR×Qual interaction, and on the assumption that the promotion-production function is invariant to the source of the rating. The ΔCS = PCS - Pactual comparisons in Table 5 treat peer-only, head-only, nonpeer-only, and self-only rating regimes as if the logit coefficients estimated on the actual mixed-rating data remain valid under each counterfactual. This is a strong extrapolation and is not discussed. Since the paper itself frames this section as exploratory, the claims about the promotion impact of strategic behavior should be correspondingly hedged, and the dependence of the simulation on the unvalidated logit specification should be acknowledged.","section":"§5.2, Table 4"}],"minor_comments":[{"comment":"The table numbering is inconsistent: the text says 'Table 2 presents the summary statistics' for the self-assessment table, but that table is labeled 'TABLE 1'; the subsequent correlation matrix is labeled 'TABLE 2' though the text refers to it as 'Table 1'. Renumber the tables or fix the cross-references.","section":"§4.1, §4.2, Tables 1 and 2"},{"comment":"The phrase 'In this session' appears in §5 and in the Related Work section; 'session' should be 'section'.","section":"§5 and Related Work"},{"comment":"The dependent variable header in Table 4 reads 'PROMITION', which is a typo for 'PROMOTION'.","section":"Table 4"},{"comment":"The text refers to 'Figure 1' illustrating the qualification premium and peer differences, but no figure appears in the manuscript. Include the figure or remove the reference.","section":"Figure 1"},{"comment":"Some references are incomplete, notably [8] (no venue/year beyond 2015) and [18] (journal name 'Data Mining and Knowledge Discovery' without volume/pages). Completing these would improve the paper's scholarly apparatus.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper uses proprietary data from a single firm, which limits external reproducibility, though the estimation is simple and the key coefficients are reported. The main concern for the editor is that the headline causal interpretation ('discriminatory generosity' as strategic manipulation) is not identified under the current design; the manuscript's own Section 4.2 acknowledges the interpretation rests on an untested benchmark assumption. If the authors can either supply additional identification tests or reframe the contribution as a rigorous descriptive finding about peer-rater behavior, the paper would be a valuable empirical contribution. As it stands, the gap between the causal language in the abstract and the identifying assumptions is too large to accept without revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth reading for the empirical pattern alone: in one firm's crowdsourced reviews, same-rank peer raters give qualified colleagues lower ratings than unqualified colleagues, relative to nonpeer raters. That interaction (β3 = -0.493, SE 0.079) is large, precisely estimated, and not something the existing top-down evaluation literature shows. It generalizes lab sabotage findings to a real workplace, which is a genuine contribution.\n\nWhat the paper does well: the regression design is transparent, the ratee/rank/department/year fixed effects absorb the obvious confounders, and the authors do not oversell the counterfactual analysis—they explicitly call it exploratory. The self-assessment result (employees would raise their own percentile by about 6.5 percentile points if their ratings decided the outcome) is a clean, reproducible descriptive finding.\n\nThe soft spot is exactly what the stress-test says. The causal reading as strategic manipulation rests on two assumptions that are asserted rather than tested: that heads and nonpeer raters are an unbiased benchmark, and that Qual does not proxy for ratee-specific performance trends. Same-rank peers are the raters most likely to observe a qualified employee's actual post-qualification performance; if that performance dips after passing the objective bar, honest peer ratings would produce the same negative interaction. Likewise, if nonpeer raters anchor on formal credentials, the benchmark is contaminated. The paper's head-rating correlation evidence is suggestive but not decisive. There are no leads/lags, placebos, or rater-pair analyses to discriminate between strategic manipulation and information asymmetry. That is the main gap. The counterfactual promotion simulations also lack standard errors, though the authors label them exploratory.\n\nThe data are proprietary and the sample is 153 employees in one firm, so external validity is limited. But the findings are internally consistent and the main estimate is robust enough to deserve scrutiny. This is not a fatal flaw; the paper's own language is careful, mostly saying the pattern is 'consistent with' strategic reporting.\n\nVerdict: I would send this to peer review. It deserves a serious referee, with the clear instruction to push on identification. The pattern is real and interesting; the 'strategic manipulation' label is the part that needs more support. For a reading group, it would spark a good discussion about what counts as evidence of bias. I would cite it, with a caveat, as field evidence of peer rating distortion.\n\nRecommendation: accept for peer review.","headline":"A real field pattern of peer rating distortion; the strategic-manipulation label is plausible but not fully identified.","tokens_in":8920,"tokens_out":2169,"would_cite":true,"duration_ms":21703,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Using five years of ratings from a Chinese professional service firm, the paper shows that peer evaluators give lower scores to qualified same-rank coworkers and higher scores to unqualified ones, a pattern it calls 'discriminatory…","keywords":["peer evaluation","crowdsourced performance assessment","strategic manipulation","discriminatory generosity","promotion","bias","fairness-aware data mining","people analytics"],"falsifier":"Re-estimate the regression after adding an independent, contemporaneous measure of each ratee's actual job performance—such as billable hours, project completions, or client feedback—as a control. The discriminatory-generosity pattern should persist among ratees with equal objective performance. A sharper test is a regression-discontinuity design around the promotion-qualification threshold: peer ratings should show a discontinuous drop as soon as the ratee passes the objective requirements, while non-peer ratings should show no such drop (or a rise). If the discontinuity is absent, the strategic interpretation fails.","tokens_in":7932,"feed_emoji":"📉","tokens_out":7944,"duration_ms":71901,"temperature":0.7,"pith_summary":"This paper tries to establish that employees systematically distort the crowdsourced performance ratings they give to their same-rank coworkers. Using HR and assessment records from a Chinese professional service firm, the authors show that a worker who has already met the objective requirements for promotion receives lower ratings from peers than from non-peers, while a worker who has not yet met those requirements receives higher ratings from peers. They call this pattern \"discriminatory generosity\" and argue it reflects rivalry: the tough ratings are aimed at competent competitors, and the generous ratings for less-eligible peers serve as a mask. The stakes are practical because the firm uses these assessments to set merit pay and decide promotions. If the pattern is real, historical peer-rating data carry a built-in anti-competitiveness bias that predictive talent analytics must account for.","feed_headline":"Peer raters punish the qualified and reward the unqualified","feed_subtitle":"Same-rank coworkers in a Chinese firm downgraded peers who met promotion criteria while boosting those who hadn't.","key_machinery":"The load-bearing device is the interaction term in a linear rating regression. The paper defines two binary variables: Peer indicates that rater and ratee hold the same hierarchical rank, and Qual indicates that the ratee has passed the firm's objective promotion requirements (a combination of attendance, academic qualifications, project experience, and tenure). Fixed effects for ratee, rank, department, and year remove stable differences, so the interaction coefficient $\\beta_3$ isolates whether the rating gap between qualified and unqualified ratees changes with peer status. Under no manipulation the qualification premium should be the same for peers and non-peers, which forces $\\beta_3 = 0$; a nonzero $\\beta_3$ is the paper's measure of strategic behavior. The paper then uses these coefficient estimates to construct qualification premiums and peer differences that display the discriminatory generosity pattern, and feeds counterfactual rating scenarios into a logistic promotion model to estimate effects on promotion probability.","core_discovery":"The central discovery is the \"discriminatory generosity\" pattern in peer evaluation. In the regression $R_{ijt} = \\beta_0 + \\beta_1 \\text{Peer}_{ijt} + \\beta_2 \\text{Qual}_{jt} + \\beta_3 (\\text{Peer}_{ijt} \\times \\text{Qual}_{jt})$ plus fixed effects, the interaction coefficient $\\beta_3$ is estimated at $-0.493$ (SE $0.079$). Because $\\beta_2 = 0.165$, a qualified ratee receives a positive qualification premium from non-peer raters ($0.165$) but a negative total premium from peer raters ($0.165 - 0.493 = -0.328$). The coefficient $\\beta_1 = 0.436$ means that among ratees who have not yet passed the objective requirements, peers get substantially higher ratings than non-peers. The authors read these two effects together as strategic: same-rank raters downgrade capable competitors and overrate less eligible peers, and the generosity toward unqualified peers hides the attack on qualified ones. A separate self-assessment analysis shows that employees' own ratings would lift their percentile rank by about 6.5 percentage points, and counterfactual promotion simulations based on the firm's promotion logit indicate that self-evaluation would raise an average employee's promotion probability by roughly 3.98 percentage points (about 7.16 percent relative to the mean).","pith_inferences":["A testable extension is to interact the manipulation effect with the intensity of promotion competition: the negative $\\beta_3$ should be larger in department-rank-year cells with more candidates per open slot, if rivalry is the driving mechanism.","A sharper identification is a regression discontinuity in Qual: since Qual is defined by objective thresholds, comparing peer ratings for ratees just above and below the thresholds would isolate the strategic response from any underlying performance difference.","If the \"mask\" interpretation is right, the overrating of unqualified peers should increase as the number of qualified peers in the same rank grows, because raters need to balance the aggregate; this can be tested with cell-level shares of qualified peers.","For practitioners, the results suggest a concrete de-biasing rule: subtract the estimated peer-by-qualification interaction from peer ratings before aggregation, or drop peer ratings from high-stakes promotion decisions."],"forward_implications":["In this firm's assessment system, peer evaluations actively lower the relative standing of coworkers who have qualified for promotion, so the crowdsourced component of the review pushes against the firm's stated promotion criteria.","Because the promotion logit shows a strong positive effect of the rating percentile, the strategic distortion in peer ratings can shift promotion outcomes; the counterfactual analysis indicates that self-assessment alone would raise an employee's promotion probability by about 3.98 percentage points on average.","The bias has a masking structure: unqualified peers are overrated by peers, so the aggregate effect on a department's rankings is not simply a uniform discount for qualified workers.","For talent analytics, the findings imply that the peer-rating inputs to prediction models are strategic variables, not noisy signals of performance, and should be treated as such when used to predict promotion or turnover.","Rating-based promotion criteria interact with peer bias: the firm's top-50-percent relative ranking requirement, combined with peer leniency toward unqualified peers, can help unqualified employees clear the promotion bar."],"supporting_citations":[{"why":"Establishes the economic importance of subjective performance measures in incentive contracts, which motivates why distortions in peer ratings matter for pay and promotion.","marker":"[1]"},{"why":"Documents centrality and leniency biases in managers' top-down performance evaluations, the prior literature that this paper extends to peer raters.","marker":"[3]"},{"why":"Provides experimental evidence that peer evaluation in tournament settings induces sabotage, supporting the interpretation of the peer interaction coefficient as strategic behavior.","marker":"[7]"},{"why":"Supplies a theory of relational contracts with subjective peer evaluations, framing the crowdsourced assessment environment the paper studies.","marker":"[9]"},{"why":"Provides the industry quote that peer-driven review systems can be gamed, which is the practical concern the paper sets out to test.","marker":"[25]"}],"fun_headline_variants":["Qualified peers penalized, unqualified favored","Peer ratings: punish the competent, reward the weak","Discriminatory generosity in peer reviews","Coworker ratings biased: downrank capable, boost others","In peer scoring, being qualified is a liability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result rests on the assumption that, once ratee, rank, department, and year are controlled for, a ratee's qualification status is unrelated to unobserved changes in her true performance during the rating period, and that the non-peer and department-head raters provide an unbiased benchmark; if qualified peers are genuinely performing worse at the time of review, or if non-peer raters simply reward credentials without observing current performance, the negative interaction coefficient would not show bias.","fun_headline_variants_meta":{"raw":{"variants":["Qualified peers penalized, unqualified favored","Peer ratings: punish the competent, reward the weak","Discriminatory generosity in peer reviews","Coworker ratings biased: downrank capable, boost others","In peer scoring, being qualified is a liability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000264,"raw_usage":{"total_tokens":1653,"prompt_tokens":1045,"completion_tokens":608,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":661,"completion_tokens_details":{"reasoning_tokens":535}},"tokens_in":661,"tokens_out":608,"duration_ms":7567,"temperature":1.0,"reasoning_tokens":535,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:04:44.276993+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate the regression after adding an independent, contemporaneous measure of each ratee's actual job performance—such as billable hours, project completions, or client feedback—as a control. The discriminatory-generosity pattern should persist among ratees with equal objective performance. A sharper test is a regression-discontinuity design around the promotion-qualification threshold: peer ratings should show a discontinuous drop as soon as the ratee passes the objective requirements, while non-peer ratings should show no such drop (or a rise). If the discontinuity is absent, the strategic interpretation fails.","supporting_citations":[{"cited_title":"discriminatory generosity","cited_arxiv_id":null,"evidence_quote":"Establishes the economic importance of subjective performance measures in incentive contracts, which motivates why distortions in peer ratings matter for pay and promotion."},{"cited_title":"punishes","cited_arxiv_id":null,"evidence_quote":"Documents centrality and leniency biases in managers' top-down performance evaluations, the prior literature that this paper extends to peer raters."}],"review_version":1}