REVIEW 2 major objections 5 minor 28 references
Discovery of Bias and Strategic Behavior in Crowdsourced Performance Assessment
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Using five years of ratings from a Chinese professional service firm, the paper shows that peer evaluators give lower scores to qualified same-rank coworkers and higher scores to unqualified ones, a pattern it calls 'discriminatory…
desk verdict A real field pattern of peer rating distortion; the strategic-manipulation label is plausible but not fully identified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is the interaction term in a linear rating regression. The paper defines two binary variables: Peer indicates that rater and ratee hold the same hierarchical rank, and Qual indicates that the ratee has passed the firm's objective promotion requirements (a combination of attendance, academic qualifications, project experience, and tenure). Fixed effects for ratee, rank, department, and year remove stable differences, so the interaction coefficient $\beta_3$ isolates whether the rating gap between qualified and unqualified ratees changes with peer status. Under no manipulation the qualification premium should be the same for peers and non-peers, which forces $\beta_3 = 0$; a nonzero $\beta_3$ is the paper's measure of strategic behavior. The paper then uses these coefficient estimates to construct qualification premiums and peer differences that display the discriminatory generosity pattern, and feeds counterfactual rating scenarios into a logistic promotion model to estimate effects on promotion probability.
What would settle it
Re-estimate the regression after adding an independent, contemporaneous measure of each ratee's actual job performance—such as billable hours, project completions, or client feedback—as a control. The discriminatory-generosity pattern should persist among ratees with equal objective performance. A sharper test is a regression-discontinuity design around the promotion-qualification threshold: peer ratings should show a discontinuous drop as soon as the ratee passes the objective requirements, while non-peer ratings should show no such drop (or a rise). If the discontinuity is absent, the strategic interpretation fails.
Extended reading notes
Core claim
The central discovery is the "discriminatory generosity" pattern in peer evaluation. In the regression $R_{ijt} = \beta_0 + \beta_1 \text{Peer}_{ijt} + \beta_2 \text{Qual}_{jt} + \beta_3 (\text{Peer}_{ijt} \times \text{Qual}_{jt})$ plus fixed effects, the interaction coefficient $\beta_3$ is estimated at $-0.493$ (SE $0.079$). Because $\beta_2 = 0.165$, a qualified ratee receives a positive qualification premium from non-peer raters ($0.165$) but a negative total premium from peer raters ($0.165 - 0.493 = -0.328$). The coefficient $\beta_1 = 0.436$ means that among ratees who have not yet passed the objective requirements, peers get substantially higher ratings than non-peers. The authors read these two effects together as strategic: same-rank raters downgrade capable competitors and overrate less eligible peers, and the generosity toward unqualified peers hides the attack on qualified ones. A separate self-assessment analysis shows that employees' own ratings would lift their percentile rank by about 6.5 percentage points, and counterfactual promotion simulations based on the firm's promotion logit indicate that self-evaluation would raise an average employee's promotion probability by roughly 3.98 percentage points (about 7.16 percent relative to the mean).
Load-bearing premise
The result rests on the assumption that, once ratee, rank, department, and year are controlled for, a ratee's qualification status is unrelated to unobserved changes in her true performance during the rating period, and that the non-peer and department-head raters provide an unbiased benchmark; if qualified peers are genuinely performing worse at the time of review, or if non-peer raters simply reward credentials without observing current performance, the negative interaction coefficient would not show bias.
Editorial extensions
If this is right
- In this firm's assessment system, peer evaluations actively lower the relative standing of coworkers who have qualified for promotion, so the crowdsourced component of the review pushes against the firm's stated promotion criteria.
- Because the promotion logit shows a strong positive effect of the rating percentile, the strategic distortion in peer ratings can shift promotion outcomes; the counterfactual analysis indicates that self-assessment alone would raise an employee's promotion probability by about 3.98 percentage points on average.
- The bias has a masking structure: unqualified peers are overrated by peers, so the aggregate effect on a department's rankings is not simply a uniform discount for qualified workers.
- For talent analytics, the findings imply that the peer-rating inputs to prediction models are strategic variables, not noisy signals of performance, and should be treated as such when used to predict promotion or turnover.
- Rating-based promotion criteria interact with peer bias: the firm's top-50-percent relative ranking requirement, combined with peer leniency toward unqualified peers, can help unqualified employees clear the promotion bar.
Reading between the lines
- A testable extension is to interact the manipulation effect with the intensity of promotion competition: the negative $\beta_3$ should be larger in department-rank-year cells with more candidates per open slot, if rivalry is the driving mechanism.
- A sharper identification is a regression discontinuity in Qual: since Qual is defined by objective thresholds, comparing peer ratings for ratees just above and below the thresholds would isolate the strategic response from any underlying performance difference.
- If the "mask" interpretation is right, the overrating of unqualified peers should increase as the number of qualified peers in the same rank grows, because raters need to balance the aggregate; this can be tested with cell-level shares of qualified peers.
- For practitioners, the results suggest a concrete de-biasing rule: subtract the estimated peer-by-qualification interaction from peer ratings before aggregation, or drop peer ratings from high-stakes promotion decisions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies bias and strategic behavior in crowdsourced performance assessment using a five-year HR archive from a Chinese professional service firm, with 7,778 rater-ratee-year observations. The main analysis estimates an OLS regression of ratings on a peer-rater indicator, an objective promotion-qualification indicator, and their interaction, together with ratee, ratee-rank, department, and year fixed effects (Eq. (1), Table 3). The key finding is a significantly negative interaction coefficient (β3 = -0.493, SE 0.079), which the authors interpret as 'discriminatory generosity': peer raters downgrade qualified peers and overrate unqualified peers, relative to nonpeer raters. The paper also documents self-assessment inflation (§4.1) and presents counterfactual simulations of promotion probabilities under alternative rating schemes (§5.2). The central claim is that the observed pattern constitutes strategic manipulation in peer evaluation.
Significance. If the causal interpretation holds, this is one of the first field-data demonstrations of strategic manipulation in crowdsourced performance reviews, and the interaction-based design using an objective qualification measure is a useful template for fairness-aware talent analytics. The paper's strengths include clean OLS estimation with detailed fixed effects, transparent reporting of coefficients and clustered standard errors, and the use of an external objective benchmark (Qual) that is estimated rather than imposed. The descriptive pattern is robust and interesting: the negative interaction is large, precisely estimated, and internally consistent with the self-assessment results and the counterfactual correlations. However, the central inference from β3 to strategic behavior rests on untested identifying assumptions; the causal label is not yet established by the evidence presented.
major comments (2)
- [§4.3, Eq. (1), Table 3] The interpretation of β3 = -0.493 as evidence of strategic manipulation requires that, conditional on ratee, rank, department, and year fixed effects, Qual is uncorrelated with ratee-specific time-varying performance and that nonpeer and department-head ratings are unbiased benchmarks. Neither assumption is tested. The negative interaction is equally consistent with peers using private information that qualified peers' current performance is lower after passing requirements, or with nonpeer raters anchoring on formal credentials without observing current performance. Section 4.2 asserts that 'interpreting department heads' ratings as a nonstrategic benchmark' supports the interpretation, but this is an assertion, not a test; the higher correlation between head and nonpeer ratings (0.827) than between head and peer ratings (0.520) does not establish that nonpeers are unbiased. The manuscript should provide additional evidence for the causal reading, such as leads/lags around the qualification event, placebo tests using future qualification status, rater-pair analyses, or a direct discussion of why the information-asymmetry interpretation is ruled out. Without such evidence, the headline claim of strategic 'discriminatory generosity' is overreach.
- [§5.2, Table 4] The counterfactual simulation of promotion probabilities rests on the logit model in Table 4, which itself includes the endogenous PR×Qual interaction, and on the assumption that the promotion-production function is invariant to the source of the rating. The ΔCS = PCS - Pactual comparisons in Table 5 treat peer-only, head-only, nonpeer-only, and self-only rating regimes as if the logit coefficients estimated on the actual mixed-rating data remain valid under each counterfactual. This is a strong extrapolation and is not discussed. Since the paper itself frames this section as exploratory, the claims about the promotion impact of strategic behavior should be correspondingly hedged, and the dependence of the simulation on the unvalidated logit specification should be acknowledged.
minor comments (5)
- [§4.1, §4.2, Tables 1 and 2] The table numbering is inconsistent: the text says 'Table 2 presents the summary statistics' for the self-assessment table, but that table is labeled 'TABLE 1'; the subsequent correlation matrix is labeled 'TABLE 2' though the text refers to it as 'Table 1'. Renumber the tables or fix the cross-references.
- [§5 and Related Work] The phrase 'In this session' appears in §5 and in the Related Work section; 'session' should be 'section'.
- [Table 4] The dependent variable header in Table 4 reads 'PROMITION', which is a typo for 'PROMOTION'.
- [Figure 1] The text refers to 'Figure 1' illustrating the qualification premium and peer differences, but no figure appears in the manuscript. Include the figure or remove the reference.
- [References] Some references are incomplete, notably [8] (no venue/year beyond 2015) and [18] (journal name 'Data Mining and Knowledge Discovery' without volume/pages). Completing these would improve the paper's scholarly apparatus.
Circularity Check
No significant circularity: the central interaction coefficient is estimated from data and no prediction is constructed from the target result.
full rationale
The paper's central claim rests on the OLS estimate of the interaction term β3 = -0.493 (SE 0.079) in Equation (1), which is fit to observed rater-ratee-year ratings. The quantity labeled 'discriminatory generosity' is an interpretation of this estimated coefficient, not a quantity that was imposed or solved for. The identification argument—that β3 = 0 in the absence of strategic manipulation—is an untestable assumption about counterfactual rater behavior, but an identification assumption is not circularity: the data could have produced β3 ≈ 0 or the opposite sign. The comparison of peer versus nonpeer and head ratings uses an external benchmark (department head ratings as nonstrategic), but the paper does not derive that benchmark from the target conclusion; it treats it as an assumption and reports supporting correlations. The counterfactual promotion simulation in Section 5.2 uses the fitted logistic promotion model from Table 4 to compute changes in promotion probabilities across rating scenarios; this is an extrapolation from a fitted model, but those simulated probabilities are not fed back into the rating regression that establishes the main finding, and the simulation is explicitly framed as exploratory. The only self-reference is the footnote stating that an earlier version circulated under a different title; no load-bearing result is imported from the authors' prior work. No fitted parameter is renamed as a prediction, no uniqueness theorem is invoked, and no known result is repackaged under new coordinates. Overall, the derivation chain is self-contained relative to its empirical inputs, and the main coefficient is genuinely estimated rather than constructed.
Assumptions & free parameters
free parameters (4)
- β3 (Peer × Qual interaction coefficient) =
-0.493 (robust SE 0.079)
- β1 (Peer coefficient) =
0.436 (robust SE 0.061)
- β2 (Qual coefficient) =
0.165 (robust SE 0.042)
- Logistic promotion model coefficients (PR, Qual, PR×Qual, License) =
6.721, 1.546, -2.572, 1.373 (Table 4, Column 1)
assumptions (3)
- domain assumption Qual is exogenous to unobserved ratee performance after conditioning on the included fixed effects.
- domain assumption Non-peer raters and department heads provide an unbiased benchmark for ratee qualification.
- domain assumption The logistic promotion model estimated on actual ratings can be applied to counterfactual rating distributions.
Cite this review
Pith. "Pith review of Discovery of Bias and Strategic Behavior in Crowdsourced Performance Assessment." pith.science (2026). https://pith.science/paper/3AKYONPX
@misc{pith2026190801718,
author = {Pith},
title = {Pith review of: Discovery of Bias and Strategic Behavior in Crowdsourced Performance Assessment},
year = {2026},
howpublished = {\url{https://pith.science/paper/3AKYONPX}},
note = {Machine review of arXiv:1908.01718}
}
read the original abstract
With the industry trend of shifting from a traditional hierarchical approach to flatter management structure, crowdsourced performance assessment gained mainstream popularity. One fundamental challenge of crowdsourced performance assessment is the risks that personal interest can introduce distortions of facts, especially when the system is used to determine merit pay or promotion. In this paper, we developed a method to identify bias and strategic behavior in crowdsourced performance assessment, using a rich dataset collected from a professional service firm in China. We find a pattern of "discriminatory generosity" on the part of peer evaluation, where raters downgrade their peer coworkers who have passed objective promotion requirements while overrating their peer coworkers who have not yet passed. This introduces two types of biases: the first aimed against more competent competitors, and the other favoring less eligible peers which can serve as a mask of the first bias. This paper also aims to bring angles of fairness-aware data mining to talent and management computing. Historical decision records, such as performance ratings, often contain subjective judgment which is prone to bias and strategic behavior. For practitioners of predictive talent analytics, it is important to investigate potential bias and strategic behavior underlying historical decision records.
Reference graph
Works this paper leans on
-
[1]
Discovery of Bias and Strategic Behavior in Crowdsourced Performance Assessment∗Yifei Huang Outreach.io Seattle, USA yifei.huang@outreach.io Matt Shum California Institute of Technology Pasadena, USA mshum@caltech.edu Xi Wu Central University of Finance and Economics Beijing, China wuxi@cufe.edu.cn Jason Zezhong Xiao Cardiff University Cardiff, UK xiao@ca...
-
[2]
We can also consider this from the perspective of peer difference. Without manipulation, we expect the peer difference to be independent of whether the peer ratee has Discovery of Bias and Strategic Behavior in Crowdsourced Performance Assessment TMC 2019, Alaska, USA 3 passed requirement or not. That is, ∆Peer(1) = ∆Peer(0), which also gives to β3 =
work page 2019
-
[3]
To sum up, the interaction terms in the regression model are important for quantifying strategic manipulation, and we expect β3 = 0, in the absence of strategic manipulation. 4 RESULTS OF DETECTING STRATEGIC BEHAVIOR 4.1 Preliminary Evidence of Strategic Behavior We start by providing some simple evidence showing that employees are indeed exhibiting self-...
work page 2019
-
[4]
We can clearly see the interaction effects between ratee qualification and peer status. Qualification premium is negative when the rater is a peer, while it is positive when the rater is a nonpeer. Peer difference is positive when the ratee failed qualification, while it becomes slightly negative when the ratee is qualified. Figure 1 Ratee Qualification P...
work page 2019
-
[5]
TABLE 5 SUMMARY STATISTICS: SIMULATED CHANGES IN PROMOTION PROBABILITY Variables # Obs. Mean Median SD % negative % zero % positive ∆head 426 0.0125 0 0.1060 38.97 28.64 32.39 ∆peer 368 0.0079 0 0.1566 37.50 24.18 38.32 ∆nonpeer 426 0.0003 0 0.0677 27.46 47.18 25.35 ∆self 402 0.0398 0 0.1398 31.84 25.37 42.79 Table 6 presents the correlation matrix of the...
work page 2019
-
[10]
The “New” Performance Management Paradigm: Capitalizing on the Unrealized Potential of 360 Degree Feedback. People and Strategy. (2013)
work page 2013
-
[12]
Relational Contracts with Subjective Peer Evaluations. The RAND Journal of Economics. 47, 1 (2016), 3–28
work page 2016
-
[13]
Algorithmic Bias: From Discrimination Discovery to Fairness-aware Data Mining. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2016)
work page 2016
Show all 28 references
-
[14]
Management Science
Sabotage in Tournaments: Evidence from a Laboratory Experiment. Management Science. 57, 4 (2011), 611–627
2011
-
[19]
Data Mining and Knowledge Discovery
Measuring discrimination in algorithmic decision making. Data Mining and Knowledge Discovery. (2017)
2017
-
[20]
The American Economic Review
Optimal Contracting with Subjective Evaluation. The American Economic Review. 93, 1 (2003), 216–240
2003
-
[23]
http://fortune.com/2013/10/08/should-performance-reviews-be-crowdsourced/
2013
-
[24]
Expert Systems with Applications
Domain Driven Data Mining in Human Resource Management: A Review of Current Research. Expert Systems with Applications. (2013)
2013
-
[26]
Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2016)
Talent Circle Detection in Job Transition Networks. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2016)
2016
-
[28]
Joint Representation Learning
Person-Job Fit: Adapting the Right Talent for the Right Job with. Joint Representation Learning. ACM Trans. Manage. Inf. Syst. (2018)
2018
-
[1979]
Journal of Accounting Research
Performance Evaluation and Directed Job Effort: Model Development and Analsis in a CPA-Firm Setting. Journal of Accounting Research. 17, 2 (1979), 436–455
1979
-
[1994]
The Quarterly Journal of Economics
Subjective Performance Measures in Optimal Incentive Contracts. The Quarterly Journal of Economics. 109, 4 (1994), 1125–1156
1994
-
[1999]
Journal of Economic Literature
The Provision of Incentives in Firms. Journal of Economic Literature. 37, 1 (1999), 7–63
1999
-
[2003]
Accounting Review
Subjectivity and the weighting of performance measures: Evidence from a balanced scorecard. Accounting Review. (2003)
2003
-
[2006]
Journal of Accounting Research
Subjective Performance Indicators and Discretionary Bonus Pools. Journal of Accounting Research. (2006)
2006
-
[2010]
The American Economic Review
Tournaments and Office Politics: Evidence from a Real Effort Experiment. The American Economic Review. 100, 1 (2010), 504–517
2010
-
[2011]
Accounting Review
The Determinants and Performance Effects of Managers’ Performance Evaluation Biases. Accounting Review. 86, 5 (2011), 1549–1575
2011
-
[2013]
Management Science
Performance Appraisals and the Impact of Forced Distribution-An Experimental Investigation. Management Science. 59, 1 (2013), 54–68
2013
-
[2014]
Predicting employee expertise for talent management in the enterprise. (2014)
2014
-
[2015]
International Journal of Management Reviews
Effectiveness of Performance Appraisal: An Integrated Framework. International Journal of Management Reviews. (2015)
2015
-
[2016]
Accounting, Organizations and Society
How Control System Design Affects Performance Evaluation Compression: The Role of Information Accuracy and Outcome Transparency. Accounting, Organizations and Society. (2016)
2016
-
[2017]
Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2017)
Prospecting the Career Development of Talents:A Survival Analysis Perspective. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2017)
2017
-
[2018]
The 1st International Workshop on Organizational Behavior and Talent Analytics (Held in conjunction with KDD’18) (2018)
Which One Will be Next? An Analysis of Talent Demission. The 1st International Workshop on Organizational Behavior and Talent Analytics (Held in conjunction with KDD’18) (2018)
2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.