REVIEW 5 major objections 4 minor 37 references
Superscoring — combining best section scores across retakes — systematically inflates observed scores by selecting favorable noise, degrades signal fidelity, and widens wealth-based gaps, while Single-Sitting preserves signal but excludes h
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 00:56 UTC pith:6IXTGWIX
load-bearing objection Solid order-statistic inflation results, but the participation/exclusion claim is undercut by the model's own zero-effort outside option. the 5 major comments →
Effort Matters in Score-Based Admissions: How Retaking and Aggregation Shape Test Scores
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the score aggregation rule itself creates a structural trade-off: Superscoring systematically inflates submitted scores through order-statistic selection over independent test-day noise, degrading signal fidelity and mechanically amplifying wealth-based disparities, while Single-Sitting preserves signal fidelity but excludes high-ability students who cannot afford simultaneous preparation. Neither rule uniformly dominates; Superscoring is preferable only when the recovered constrained high-ability population is large enough to offset the noise-induced bias. The paper proves these effects in a strategic model where students allocate effort across two subjects and up
What carries the argument
The central object is the Superscoring operator — the per-section maximum across attempts — acting on a two-stage strategic effort-allocation problem with wealth-dependent costs. The load-bearing mechanism is order-statistic selection over independent noise draws: because noise is independent across sections and attempts, taking section maxima decouples favorable noise from different days, inflating expected scores and degrading the signal. This produces the Safety Net Effect (the retake option lowers first-round effort) and Effort Specialization (students stop studying subjects whose scores are already banked), which together generate the Access Premium — the expected score gap between a re
Load-bearing premise
The simulation results hinge on the calibration in Appendix F.3 that attributes only 40% of the empirical Q5-Q1 SAT gap to pre-existing latent ability and forces the remaining 60% to be generated endogenously by wealth-dependent costs, with a flat +45-point retake adjustment; if the true pre-existing share differs, the headline false-negative and composition numbers change.
What would settle it
Estimate the share of the Q5-Q1 score gap that is already present in pre-preparation ability measures (e.g., PSAT or 8th-grade scores); if that share is substantially above the 40% assumed in the calibration, the simulated amplification results in Tables 1 and 2 would not survive, whereas if retake score gains match the order-statistic noise premium, the inflation mechanism is confirmed.
If this is right
- Universities using Superscoring admit a class whose observed scores overstate latent ability, because expected scores rise with more attempts even when total study effort falls.
- Wealthier applicants receive a larger mechanical score boost from repeated testing, so the submitted-score gap between high- and low-wealth groups is strictly wider under Superscoring than under Single-Sitting.
- Single-Sitting keeps scores honest but pushes high-ability, resource-constrained students out of the applicant pool entirely, because the simultaneous preparation burden is superadditive.
- FISC and the Variance Tax can neutralize noise-harvesting inflation without eliminating the participation benefit of Superscoring, but they are demographic-blind and hit a structural ceiling on representation.
- The Contextual Correction equalizes false-negative rates across wealth groups and is the only tested mechanism that brings admitted-class composition close to the true ability distribution.
- There is no uniformly dominant scoring rule: the choice between Superscoring and Single-Sitting is a genuine precision-versus-fairness trade-off, and the preferred rule depends on how large the pool of recovered high-ability constrained students is.
Where Pith is reading between the lines
- The order-statistic inflation mechanism generalizes beyond the SAT: any selection process that takes the best of several noisy evaluations — repeated job interviews, portfolio reviews, or benchmark runs — will overstate typical performance and reward whoever can afford more attempts.
- A testable extension is to compare first-attempt and best-of-many scores for the same cohort; if the retake gain matches the order-statistic noise premium rather than genuine learning, the inflation mechanism is confirmed, and if it exceeds it, some learning is being mislabeled as noise.
- The calibration choice that attributes 40% of the empirical wealth-score gap to pre-existing ability is the most sensitive input; re-estimating that share from longitudinal pre-test measurements would materially change the headline false-negative and composition numbers.
- The Contextual Correction result implies that institutions seeking ability-proportionate admissions may need to accept explicit demographic adjustments, since measurement-only fixes hit a structural ceiling — a policy question the paper notes as a disparate-treatment concern.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a two-stage strategic model in which students with wealth-dependent effort costs choose preparation and retake effort in response to an admissions rule, and compares Single-Sitting (best complete attempt) with Superscoring (best section scores). It claims that Superscoring inflates scores through order-statistic selection over noise, degrades signal fidelity, and amplifies wealth disparities, while Single-Sitting preserves signal fidelity but excludes high-ability constrained students. It then proposes three algorithmic corrections (FISC, Variance Tax, Contextual Correction) and evaluates them in simulations calibrated to 2025 College Board SAT data.
Significance. The paper addresses a timely and policy-relevant mechanism: score aggregation rules are not passive measurements but shape student effort and retake incentives. Proposition 4, the order-statistic inflation result, is correct and cleanly proved. The paper also makes a useful modeling contribution by embedding wealth frictions in a strategic test-taking framework, and the reproducibility measures (seeded simulations, planned code release, uniform-population baselines, parameter-grid robustness) are strengths. If the central trade-off were fully established, the paper would be a valuable contribution to the strategic-classification and admissions-design literatures. However, as detailed below, several load-bearing claims are not supported as stated, particularly the participation-exclusion result and the signal-degradation claim in the convex-cost regime used by the simulations.
major comments (5)
- [Proposition 8 and Appendix D.7] The Participation Effect is internally inconsistent with the model. Because f(0)=0 and costs are zero at zero effort, choosing x1=y1=0 and not retaking yields expected utility 0 under both rules, so V_1^SS >= 0 for every student. No student can have V_1^SS < 0, contradicting the proposition's claimed nonempty set. The proof in D.7 also has a logical gap: it shows V_SC_split > V_SS(xbar,ybar) and then infers V_SC* > V_SS*, but V_SS* >= V_SS(xbar,ybar), so the chain is invalid. Since Proposition 9 and the abstract's claim that Single-Sitting 'excludes high-ability constrained students' depend on Proposition 8, the precision-fairness trade-off is unsupported as stated.
- [Proposition 5 and Appendix C] The 'signal degradation' claim is not established in the convex-cost regime used for simulations. Section 5.1 presents Proposition 5 as showing total expected effort falls under Superscoring, but Appendix C explicitly states that the total-effort-reduction component 'does not generalize cleanly to strict convexity' and that the net direction is parameter-dependent when alpha>1. Appendix F.2 uses alpha=2. Consequently, the main text's claim that Superscoring 'degrades signal accuracy' through lower human-capital investment is unsupported in the empirical setting, and the simulations cannot be cited as evidence for Proposition 5.
- [Sections 6.1, 6.2, Appendix F.2] FISC is described as a structural scoring rule that changes student effort and retake behavior ('by taxing multiple attempts, it deters students from retaking'), but the simulation treats FISC as an ex-post score transformation applied to the Superscoring equilibrium. Appendix F.2 states that FISC and CTX are 'evaluated as ex-post score transformations applied to the equilibrium effort allocations and retake decisions generated under the Superscoring regime.' Therefore the Table 1 FISC results (e.g., Q1 FNR 52%) do not measure the mechanism claimed in Section 6.1, and the statement that FISC 'successfully deters speculative retaking' is not supported by the experiments.
- [Appendix F.3 and Tables 1-2] The simulation's wealth-disparity results are partly a restatement of the calibration target. Appendix F.3 chooses a 40% baseline-ability discount precisely so that 60% of the empirical Q5-Q1 SAT gap is generated endogenously by the model's wealth-dependent costs, plus a flat +45-point retake adjustment. Tables 1-2 then report large wealth-based FNR gaps and composition distortions. These quantitative findings are conditional on an unvalidated decomposition of the observed gap into latent ability versus endogenous effort; an independent moment or out-of-sample validation is needed before the simulation can be read as confirming the theoretical amplification result.
- [Proposition 6 and Appendix D.4] The Access Premium result is stated unconditionally for linear costs, but the proof requires 'sufficiently dispersed noise' and invokes Proposition 5, which is itself conditional. The mimicking strategy shows only that Student A can feasibly achieve a higher expected score than Student B; it does not show that the utility-maximizing choice does so, since A may trade score for effort savings. Without a formal derivation that the optimal effort choice preserves the strict score gap, Proposition 6 and the subsequent 'Access Premium' language are not rigorously established as stated.
minor comments (4)
- [Section 5.1, Proposition 5 statement] The statement says total effort falls 'whenever the expected mechanical score boost exceeds the expected score loss from reduced effort,' but the displayed inequality is about E[e_SC] < E[e_SS]. Clarify the logical relation between the condition and the two inequalities.
- [Table 11 and Appendix H.1] The table contains typos: '0.7&' should be '0.7%', and the title reads 'False Postive Rate'.
- [Figure 1 caption] The caption mixes references to SS and SC ('blue, with SS on the left and SC on the right') in a way that is difficult to parse; please redraw or clarify which panel corresponds to which rule.
- [Appendix D.7] The proof introduces 'Starget' and uses E[S_SS] >= S_target, but the inequality direction is stated without explanation; since noise can lower the realized composite, the expectation of the max is indeed at least the deterministic score, but this should be stated explicitly.
Circularity Check
No load-bearing circularity: the theoretical results are derived from model primitives, and the simulation calibration is transparent and not presented as an independent prediction.
full rationale
The central theoretical claims—score inflation under Superscoring, the Access Premium, and the wealth-gap comparison between rules—are derived from explicit model primitives and order-statistic inequalities (Propositions 4–7, 9). They do not assume the conclusion or a fitted value as an input. Self-citations appear only as related-work framing, not as load-bearing justification for a theorem. The simulation calibration in Appendix F.3 does fit a structural discount (40% of the Q5–Q1 gap assigned to latent ability) and a +45-point retake adjustment to reproduce the 2025 College Board wealth gap, but the paper labels this as calibration rather than prediction; the headline FNR and composition results are model outputs that are not identical to the calibration target. Thus there is no fitted-input-called-prediction circularity in the sense required. Separately, Proposition 8's participation/exclusion claim has a logical gap—zero effort yields utility 0, so V_SS can never be negative; and the proof compares a feasible split strategy with a non-optimal SS strategy. This is a correctness/support defect in the derivation chain, not a circularity, so it does not raise the circularity score.
Axiom & Free-Parameter Ledger
free parameters (5)
- Baseline ability discount =
40%
- Retake adjustment =
+45 score points
- Score noise volatility sigma =
≈0.2783
- Wealth gradients for latent ability =
Math 134 pts, Verbal 129 pts (empirical gaps, discounted 40%)
- Marginal effort cost schedule =
k_i = 0.14 - 0.02 w_i
axioms (7)
- domain assumption Variable effort cost has power-law form c(e|W)=k(W)e^alpha with k strictly decreasing in wealth (Assumption 1).
- domain assumption Score production f is strictly increasing and concave with f(0)=0 (Assumption 2).
- domain assumption Test-day noise is mean-zero, symmetric, continuous, with Var(delta)=sigma^2 I, independent across subjects and attempts.
- ad hoc to paper Assumption 3: beta k/(w_j eta_j) <= f'(0), guaranteeing positive optimal effort.
- ad hoc to paper Stage 1 objective is strictly concave under 'sufficiently dispersed noise' (Lemma 2, App B.2).
- domain assumption Students are expected-utility maximizers with full rationality and full knowledge of the scoring rule.
- domain assumption In Prop 9, Type B noise harvesters have lower ability than Type A and log-concave noise satisfies MLRP.
read the original abstract
Observed standardized test scores are the result of an endogenous process: students strategically allocate effort across multiple retake attempts to improve their outcomes. Because students differ in their ability to make these investments, the interaction between applicant strategy and institutional scoring rules---such as the widely used Single-Sitting and Superscoring policies---can disparately distort observed scores. We develop a strategic framework where students allocate effort in response to different scoring policies. We show that Superscoring---the practice of combining the best section scores across attempts---introduces systematic score inflation through order-statistic selection over noise draws. This degrades signal accuracy and amplifies wealth-based disparities by disproportionately rewarding applicants who can afford repeated testing. Conversely, Single-Sitting---which keeps the best overall score rather than section-level scores---preserves signal fidelity but excludes high-ability students who lack the resources to prepare for all subjects simultaneously. Neither rule uniformly dominates; instead, they force a structural trade-off between statistical precision and fair outcomes. Finally, to address this, we propose three algorithmic interventions which either modify how scores from multiple attempts are combined, or apply a post-hoc correction to observed scores. Using simulations calibrated to 2025 College Board data, we compare standard scoring rules against these proposed interventions.
Figures
Reference graph
Works this paper leans on
-
[1]
2025 Total Group SAT Suite of Assessments Annual Report , author =
2025
-
[2]
The Quarterly Journal of Economics , volume =
Chetty, Raj and Hendren, Nathaniel and Kline, Patrick and Saez, Emmanuel , title =. The Quarterly Journal of Economics , volume =. 2014 , month =. doi:10.1093/qje/qju022 , url =
-
[3]
The Quarterly Journal of Economics , volume =
Chetty, Raj and Deming, David J and Friedman, John N , title =. The Quarterly Journal of Economics , volume =. 2026 , month =. doi:10.1093/qje/qjaf050 , url =
-
[4]
The college application gauntlet: A systematic analysis of the steps to four-year college enrollment
Klasik, Daniel. The college application gauntlet: A systematic analysis of the steps to four-year college enrollment. Res. High. Educ
-
[5]
American Economic Journal: Applied Economics , Volume =
Bulman, George , Title =. American Economic Journal: Applied Economics , Volume =. 2015 , Month =. doi:10.1257/app.20140062 , URL =
-
[6]
One-Offs
Hoxby, Caroline and Avery, Christopher , year =. The Missing "One-Offs": The Hidden Supply of High-Achieving, Low-Income Students , volume =. Brookings Papers on Economic Activity , doi =
-
[7]
Sackett and Nathan R
Paul R. Sackett and Nathan R. Kuncel and Adam S. Beatty and Jana L. Rigdon and Winny Shen and Thomas B. Kiger , journal =. The Role of Socioeconomic Status in SAT-Grade Relationships and in College Admissions Decisions , urldate =
-
[8]
and Clotfelter, Charles T
Vigdor, Jacob L. and Clotfelter, Charles T. , title =. 2003 , doi =. https://jhr.uwpress.org/content/38/1/1.full.pdf , journal =
2003
-
[9]
Assessment in American Higher Education: The Role of Admissions Tests , volume =
Zwick, Rebecca , year =. Assessment in American Higher Education: The Role of Admissions Tests , volume =. The ANNALS of the American Academy of Political and Social Science , doi =
-
[10]
Dropping standardized testing for admissions trades off information and access
Garg, Nikhil and Li, Hannah and Monachou, Faidra. Dropping standardized testing for admissions trades off information and access. Manage. Sci
-
[11]
2021 , eprint=
Test-optional Policies: Overcoming Strategic Behavior and Informational Gaps , author=. 2021 , eprint=
2021
-
[12]
2024 , eprint=
Test-Optional Admissions , author=. 2024 , eprint=
2024
-
[13]
Journal of the European Economic Association , volume =
Frankel, Alex and Kartik, Navin , title =. Journal of the European Economic Association , volume =. 2022 , month =. doi:10.1093/jeea/jvab017 , url =
-
[14]
The Review of Economic Studies , volume =
Krishna, Kala and Lychagin, Sergey and Olszewski, Wojciech and Siegel, Ron and Tergiman, Chloe , title =. The Review of Economic Studies , volume =. 2026 , month =. doi:10.1093/restud/rdaf033 , url =
-
[15]
2026 , eprint=
The Impact of Competition on Outcomes of Score-Based College Admissions , author=. 2026 , eprint=
2026
-
[17]
The Quarterly Journal of Economics , volume =
Kleinberg, Jon and Lakkaraju, Himabindu and Leskovec, Jure and Ludwig, Jens and Mullainathan, Sendhil , title =. The Quarterly Journal of Economics , volume =. 2018 , month =. doi:10.1093/qje/qjx032 , url =
-
[18]
Available at SSRN 6248078 , year=
Admission Policies with Exam Retakes: Efficiency-Equity Trade-offs in Multidimensional Scoring , author=. Available at SSRN 6248078 , year=
-
[19]
2016 , eprint=
Inherent Trade-Offs in the Fair Determination of Risk Scores , author=. 2016 , eprint=
2016
-
[20]
All: Equity and Accuracy of Standardized Test Score Reporting , author=
Best vs. All: Equity and Accuracy of Standardized Test Score Reporting , author=. 2021 , eprint=
2021
-
[21]
, title =
Borghesan, E. , title =
-
[22]
2019 , eprint=
The Disparate Equilibria of Algorithmic Decision Making when Individuals Invest Rationally , author=. 2019 , eprint=
2019
-
[23]
AJ Alvero and Sonia Giebel and Ben Gebre-Medhin and anthony lising antonio and Mitchell L. Stevens and Benjamin W. Domingue , title =. Science Advances , volume =. 2021 , doi =. https://www.science.org/doi/pdf/10.1126/sciadv.abi9031 , abstract =
-
[24]
2018 , eprint=
Access to Population-Level Signaling as a Source of Inequality , author=. 2018 , eprint=
2018
-
[25]
Jesse M. Rothstein , keywords =. College performance predictions and the SAT , journal =. 2004 , note =. doi:https://doi.org/10.1016/j.jeconom.2003.10.003 , url =
-
[26]
Who benefits from SAT prep?: An examination of high school context and race/ethnicity
Park, Julie J and Becks, Ann H. Who benefits from SAT prep?: An examination of high school context and race/ethnicity. Rev. High. Ed
-
[27]
Proceedings of the 2016 ACM Conference on Innovations in Theoretical Computer Science , pages =
Hardt, Moritz and Megiddo, Nimrod and Papadimitriou, Christos and Wootters, Mary , title =. Proceedings of the 2016 ACM Conference on Innovations in Theoretical Computer Science , pages =. 2016 , isbn =. doi:10.1145/2840728.2840730 , abstract =
arXiv 2016
-
[28]
2020 , eprint=
The Role of Randomness and Noise in Strategic Classification , author=. 2020 , eprint=
2020
-
[29]
Proceedings of the 2018 ACM Conference on Economics and Computation , pages =
Dong, Jinshuo and Roth, Aaron and Schutzman, Zachary and Waggoner, Bo and Wu, Zhiwei Steven , title =. Proceedings of the 2018 ACM Conference on Economics and Computation , pages =. 2018 , isbn =. doi:10.1145/3219166.3219193 , abstract =
arXiv 2018
-
[30]
Proceedings of the 2019 ACM Conference on Economics and Computation , pages =
Kleinberg, Jon and Raghavan, Manish , title =. Proceedings of the 2019 ACM Conference on Economics and Computation , pages =. 2019 , isbn =. doi:10.1145/3328526.3329584 , abstract =
arXiv 2019
-
[31]
2021 , eprint=
Stateful Strategic Regression , author=. 2021 , eprint=
2021
-
[32]
2021 , eprint=
Gaming Helps! Learning from Strategic Interactions in Natural Dynamics , author=. 2021 , eprint=
2021
-
[33]
Proceedings of The 25th International Conference on Artificial Intelligence and Statistics , pages =
Strategic ranking , author =. Proceedings of The 25th International Conference on Artificial Intelligence and Statistics , pages =. 2022 , editor =
2022
-
[34]
Milli, Smitha and Miller, John and Dragan, Anca D. and Hardt, Moritz , title =. Proceedings of the Conference on Fairness, Accountability, and Transparency , pages =. 2019 , isbn =. doi:10.1145/3287560.3287576 , abstract =
arXiv 2019
-
[35]
Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages =
Estornell, Andrew and Das, Sanmay and Liu, Yang and Vorobeychik, Yevgeniy , title =. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2023 , isbn =. doi:10.1145/3593013.3594006 , abstract =
arXiv 2023
-
[36]
Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =
Garg, Nikhil and Li, Hannah and Monachou, Faidra , title =. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2021 , isbn =. doi:10.1145/3442188.3445889 , abstract =
arXiv 2021
-
[37]
Advances in Neural Information Processing Systems , volume=
Incentivizing desirable effort profiles in strategic classification: The role of causality and uncertainty , author=. Advances in Neural Information Processing Systems , volume=
-
[38]
Review of economic design , volume=
The theory of contests: a survey , author=. Review of economic design , volume=. 2007 , publisher=
2007
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.