REVIEW 4 major objections 4 minor 19 references
Prestige in Numbers: How Test Scores and Choices Reveal School Rankings
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read From where applicants send their GMAT scores — and nothing else — this paper recovers a ranking of U.S. MBA programs that matches the leading published ranking (correlation 0.72 on the top 100) using no school-supplied data.
desk verdict A promising data source and a plausible revealed-preference idea, but the empirical ranking method is not formally identified from the theory, and the proof has a technical gap that needs fixing before this is fully convincing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the score-monotonicity of the application distribution. In the model each school $j$ has an admission threshold $t_j$ and admits an applicant with score $s_i$ when $s_i + \varepsilon_{i,j} \geq t_j$, so the acceptance probability is $F(s_i - t_j)$; with $F$ concave, the upper-tail distribution $\bar G_s(c)$ of schools chosen by score-$s$ applicants rises in $s$ in the sense of first-order stochastic dominance. The invertible consequence is that a school's application share $g_s(c)$ should be increasing in score for schools near the top of the order, and the estimators exploit exactly this: the $m$-measure accumulates $(g_{s'}(c) - g_s(c))(s' - s)$ over all score pairs, the $m^+$ version counts only the sign-correct comparisons, and the tournament scores a school whenever its selection by a higher-scoring applicant is not echoed by a lower-scoring one.
What would settle it
Estimate the admission-error distribution $F$ from real admissions decisions: if its density rises anywhere over the relevant score range, $F$ is not concave, and the guarantee that higher scores mean more aggressive portfolios fails. A direct test of the empirical claim is to compute $\bar G_s(c)$ from the score-sending data and check whether the distribution of chosen schools is first-order stochastically increasing in score for every pair of scores; a single persistent reversal — some score level whose applicants systematically aim at less selective schools than lower-scoring applicants do — would refute the monotonicity on which both estimators rest.
Extended reading notes
Core claim
The paper's central claim is that an ordering of schools by prestige is identifiable from application portfolios and standardized test scores alone. The theoretical engine is Theorem 1: when admission requires a score plus an idiosyncratic shock with a concave distribution $F$, and when application budgets are weakly increasing in score, the optimal portfolio of a higher-scoring student first-order stochastically dominates that of a lower-scoring student — higher scorers apply to more selective schools, position by position. Because selectivity is taken to align with perceived quality, the ordering the method targets is precisely the one that would make every school's application share monotone in score, and the paper recovers it with two estimators: an $m$-measure that grades each school by the monotonicity of its application-share curve in score, and a tournament that credits a higher-scoring applicant's school choice whenever it diverges from a lower-scoring applicant's. On nine years of GMAT score-sending data both estimators reproduce the familiar hierarchy of elite MBA programs, correlating 0.72 with the leading published top-100 ranking.
Load-bearing premise
The argument depends on the idiosyncratic part of admissions being spread by a concave probability distribution; with a bell-shaped (normal or logistic) error distribution, which has an increasing density over part of its range, higher-scoring students are not guaranteed to apply to more selective schools, and the two ranking estimators lose their theoretical foundation.
Editorial extensions
If this is right
- Because the method needs only the score reports the test administrator already holds, every program with enough reports can be ranked, not merely the top 50 or 100 that expert rankings cover.
- The same machinery yields custom rankings for subgroups — undergraduate major, citizenship, or time period — with rank correlations of 0.93, 0.71, and 0.86 across those splits, revealing sensible differences such as international students favoring schools in large cities.
- A ranking built from test-administrator data is insulated from the documented manipulation of institution-supplied rankings: schools cannot adjust the inputs the way they can report class sizes, faculty counts, or peer ratings.
- The theory predicts bunching of test scores near admission thresholds and positive assortative matching of scores to school selectivity, and both patterns appear in the data.
- Keeping only the best test attempt for repeat test-takers leaves the ranking essentially unchanged (correlation 0.98 with the full-sample ranking), indicating the result is not driven by retake behavior.
Reading between the lines
- The same revealed-preference logic should transfer to any standardized test whose score-sending or score-reporting records exist — SAT, GRE, LSAT, MCAT — so undergraduate, law, and medical programs could be ranked from decentralized applicant behavior alone; this is an extension the paper does not run.
- The concavity assumption puts a boundary on where the theory bites: if admission errors are near-normally distributed, the guarantee is most secure where the error density is falling and least secure in its increasing-density lower tail, so the empirical rankings are on firmest ground for the high-scoring applicants who dominate selective schools.
- An implicit consequence is that the recovered order measures the applicant pool's perceived prestige at application time, not realized post-graduation value; a school with an inflated reputation would be ranked by belief rather than outcome, so the method is best read as an aggregate belief measure.
- The subgroup rankings open the door to on-demand personalized rankings by any observable trait recorded at test registration — intended industry, geography, or undergraduate field — which expert surveys could never scale to.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a revealed-preference method for ranking colleges and professional schools using standardized-test score-sending data. A theoretical model of application portfolios yields Theorem 1: if the admission shock distribution F is concave and application budgets are weakly increasing in score, then higher-scoring students' chosen schools first-order stochastically dominate lower-scoring students' choices. The authors then implement two estimators on GMAT score reports for U.S. full-time MBA programs (over 490,000 candidates, 688 programs): the m-measure, which ranks schools by a score-weighted monotonicity statistic of application shares, and a pairwise tournament. The resulting rankings are compared with U.S. News & World Report (USNWR) rankings, with Spearman correlations of 0.92/0.94 for the top 50 and 0.72 for the top 100.
Significance. If the central identification step were established, the paper would offer a genuinely useful alternative to institution-supplied rankings: it uses comprehensive administrative data, is resistant to school manipulation, and can be customized to student subgroups. The paper also connects the portfolio-choice literature (Chade-Smith, Ali-Shorrer) to a concrete ranking application. However, the gap between Theorem 1 and the estimators is load-bearing: the paper does not prove that the m-measure or tournament recovers the model's selectivity order. The reported correlations with USNWR are suggestive but may reflect a mechanical relation with applicant score averages rather than the revealed-preference mechanism. The model itself also has a technical gap in the proof of Lemma 3 under the stated concavity assumption.
major comments (4)
- [Section 2.2, Lemma 3] The proof of Lemma 3 asserts that s ↦ F(s-t_k) - F(s-t̃_j) is strictly monotone increasing 'as F is concave and t_k > t̃_j'. This is incorrect: concavity of F implies the difference is nondecreasing, not strictly increasing; strict monotonicity requires strict concavity (a strictly decreasing density). For a linear F, which is concave, the difference is constant. The strict monotonicity is used in the revealed-preference step to conclude that the higher-score student strictly prefers the more selective school, so Theorem 1 is not proven under the stated assumptions. Please either strengthen the assumption to strict concavity or supply a different proof, and discuss whether common choices such as normal or logistic F, which are not concave over their full support, satisfy the condition.
- [Section 3.2.1, Eq. for m1(c)] The paper asserts, without proof, that ranking schools by m1(c) = Σ_{s,s'}(g_{s'}(c)-g_s(c))(s'-s) recovers the threshold order ≥. Theorem 1 only establishes FOSD for the upper-tail distributions Ḡ_s(c) = Pr(chosen school ≥ c); it does not imply that the individual shares g_s(c) are monotone in s for each school, nor that the weighted sum m1(c) is monotone in rank. The recursive argument in Section 3.2.1 shows only that the second-ranked school is the one with the second-highest m1 value, because the added term is independent of c. This is a missing identification step connecting the theory to the main estimator. The external correlation with USNWR does not fill the gap, because m1 is a score-weighted average of application propensities and is likely collinear with average applicant GMAT score, an input to USNWR. Please provide a formal statement, with explicit assumptions, under which the m1 order recovers ≥, or characterize what the m1 order identifies.
- [Section 3.2.2] The tournament method has the same identification gap. The rule that a higher-score student who chose c' and not c gives a point for c' over c is motivated by Lemma 3, but Lemma 3 concerns a fixed utility function and an ordered portfolio of schools. The empirical tournament aggregates across heterogeneous utility functions, portfolio sizes, and score levels; no theorem shows that the resulting win counts define a transitive order or agree with ≥. The reported rank correlations with USNWR are not a substitute for this justification. Please either prove that the tournament recovers the true order under the model, or explicitly label it as a heuristic and discuss what properties it does guarantee.
- [Section 4.1 and 4.4] The validation correlations (ρ=0.92 and 0.94 for the top 50, ρ=0.72 for the top 100) are against USNWR, which itself uses the average GMAT scores of enrolled students as an input. To support the revealed-preference interpretation, the paper should compare its rankings with a simple ranking by average applicant GMAT score or by total number of score reports, and show that the m-measure and tournament add information beyond these 'mechanical' correlates. Without such a baseline, the correlation may reflect school selectivity rather than the model's comparative-statics mechanism.
minor comments (4)
- [Section 3.2.1] The displayed formula for m+(c) contains ambiguous notation: the indicator function is applied to the whole product, but the text reads as if the indicator applies only to the sign of the difference. Please clarify.
- [Table 5] Table 5 lists 'University of North Carolina - Chapel Hill' twice (rows 20 and 35) without an asterisk or footnote explaining whether these are different programs (e.g., the regular two-year program versus a one-year or global program). Please annotate the table accordingly.
- [Section 2.1] The score set S is defined with s1 > s2 > ... > sT, so statements like 's ≤ s′ implies k_v_s ≤ k_v_s′' are confusing because the ordering of indices is opposite to the numeric score ordering. Consider using an ascending score ordering or explicit score labels to reduce ambiguity.
- [Section 4.3] The sentence 'The Spearman’s rank correlation coefficient for the different rank orderings in Table 3 is 0.93' is imprecise: Table 3 contains two rankings (business/economics majors and other majors), so there is only one correlation coefficient between the two, not 'different rank orderings' in plural.
Circularity Check
No definitional or fitted-input circularity: Theorem 1 is a genuine comparative static and neither estimator is fit to the U.S. News benchmark. Residual circularity: the m-measure is by construction a monotone transform of applicant GMAT averages, while U.S.
-
renaming known result
[Section 3.2.1 (m(c) definition) and Section 2.1 (GMAT in USNWR metrics); results in Sections 4.1 and 6]
"Test scores are particularly subject to selectivity, given that the average scores of accepted students feature in the metrics that expert opinion sources use to rank schools (e.g. US News and World Report). ... m(c) = Σ_{s,s′∈S}(g_{s′}(c) − g_s(c))(s′ − s)"
Expanding the paper's own formula gives m(c) = 2T·A(c)·(E_{1/N_s}[s|c] − mean_s): by construction a positive multiple of the reweighted average GMAT score of c's applicants. U.S. News inputs average GMAT scores of enrolled students (the paper concedes test scores 'feature in the metrics'), and enrollees come from the same applicant pool. So the headline validations (ρ = 0.72 for the top 100; 0.92/0.94 for the top 50) are partly mechanical: they largely certify that GMAT-sorted applicant pools match GMAT-sorted enrolled pools, driven by the shared test-score signal rather than by the revealed-preference mechanism. The external benchmark is not fully independent of the estimator's main input, so part of the claimed validation is built into the choice of measure.
-
other
[Section 3.2.1, identification conclusion of the recursive m-measure argument]
"Given that only the first term, Σ_{s,s′∈S}(g_{s′}(c) − g_s(c))(s′ − s), depends on c ̸= c1, this school is the one yielding the second-highest value of m1(c). Therefore, the ranking produced by m1 yields an estimate of the ranking of schools that is captured by ≥."
The paper's own decomposition shows the recursive 'upper-tail' step adds only a constant independent of c (m2(c) = m1(c) + m1(c1)), so the produced ranking is, by construction, exactly descending m1 values; the upper-tail FOSD property of Theorem 1 — the only theoretical input — is never evaluated for intermediate ranks. The conclusion that this order 'is captured by ≥' is the unproven converse of the paper's argument, which established only that under the true ranking the tail distributions are monotone in score, not that the measure-maximizer is the next school in the true order. Hence the recovered ranking is stipulated to equal the model's ranking rather than derived from it; the central empirical 'prediction' coincides with the estimator's definition plus an unverified converse.
full rationale
The derivation chain's core step, Theorem 1, is genuinely self-contained: the concavity of F and monotone budgets are stated assumptions, Lemma 3 is proved via a Milgrom–Shannon increasing-differences argument, and the conclusion (upper-tail FOSD by score) is a non-trivial comparative static, not an identity. No load-bearing self-citation exists — the paper's citations (Chade–Smith, Ali–Shorrer, Avery et al., Milgrom–Shannon) are all external, and none of the central premises rests on the authors' own prior work. Neither estimator is fitted to the benchmark: m1, m+ and the tournament are closed-form functions of the application data, and the U.S. News comparison is an external, unfitted check that could have failed; this alone keeps the paper well below the high end of the circularity scale. The residual circularity is of two kinds. First (mechanical validation): expanding the m-measure gives a positive multiple of the excess reweighted-average GMAT score of a school's applicants, while U.S. News ranks MBA programs partly from average GMAT scores of enrolled students; the paper itself notes test scores 'feature in the metrics that expert opinion sources use to rank schools,' so the 0.72–0.94 validation correlations partly certify the near-tautology that high-GMAT applicant pools match high-GMAT enrolled pools. Second (stipulated identification): Section 3.2.1 asserts that descending m1 yields the ranking 'captured by ≥'; the paper's own recursion equation (m2 = m1 + constant) shows the procedure never invokes upper-tail FOSD at intermediate ranks, and the maximizer-is-next-school claim is the unproven converse of the proven implication. I flag this as an omitted identification proof and weigh it as a validity risk rather than full circularity, because the benchmark is external and unfitted and the theoretical core is independent. The proportionate score is 2: the derivation is not circular, but a meaningful mechanical component and a stipulated identification step remain.
Assumptions & free parameters
free parameters (2)
- minimum score report threshold =
122
- period cut for early/later ranking =
July 1, 2011
assumptions (5)
- domain assumption F is absolutely continuous and concave
- domain assumption Score reports are faithful proxies for applications (R^2=0.98 per GMAC)
- domain assumption Admission decisions are independent across schools and based on si+epsilon_i,j >= t_j
- domain assumption Higher scorers have at least as many application slots (kv_s <= kv_s')
- ad hoc to paper Duplicate schools exist
Cite this review
Pith. "Pith review of Prestige in Numbers: How Test Scores and Choices Reveal School Rankings." pith.science (2026). https://pith.science/paper/6MEPQDNU
@misc{pith2026250521063,
author = {Pith},
title = {Pith review of: Prestige in Numbers: How Test Scores and Choices Reveal School Rankings},
year = {2026},
howpublished = {\url{https://pith.science/paper/6MEPQDNU}},
note = {Machine review of arXiv:2505.21063}
}
abstract
This paper introduces a novel revealed-preference approach to ranking colleges and professional schools based on applicants' choices and standardized test scores. Unlike traditional rankings that rely on data supplied by institutions or expert opinions, our methodology leverages the decentralized beliefs of potential students, as revealed through their application decisions. We develop a theoretical model where students with higher test scores apply to more selective institutions, allowing us to establish a clear relationship between test score distributions and school prestige. Using comprehensive data from over 490,000 GMAT test-takers applying to U.S. full-time MBA programs, we implement two ranking methods: one based on monotone functions of test scores across schools, and another using score-adjusted tournaments between school pairs. Our approach has distinct advantages over traditional rankings: it reflects the collective judgment of the entire applicant pool rather than a small group of experts, and it utilizes data from an independent testing organization, making it resistant to manipulation by institutions. The resulting rankings correlate strongly with leading published MBA rankings ($\rho = 0.72$) while offering the additional benefit of being customizable for different student subgroups. This method provides a transparent alternative to existing ranking systems that have been subject to well-documented manipulation.
Figures
Reference graph
Works this paper leans on
-
[1]
American Economic Review, 115, 571--598
Ali, S Nageeb and Ran I Shorrer (2025), Hedging When Applying: Simultaneous Search with Correlation . American Economic Review, 115, 571--598. alishorrer24
work page 2025
-
[2]
Handbook of the Economics of Education, 7, 1--60
Angrist, Joshua, Peter Hull, and Christopher Walters (2023), Methods for Measuring School Effectiveness . Handbook of the Economics of Education, 7, 1--60. angrist2023methods
work page 2023
-
[3]
Avery, Christopher N., Mark E. Glickman, Caroline M. Hoxby, and Andrew Metrick (2012), A Revealed Preference Ranking of U.S. Colleges and Universities . Quarterly Journal of Economics, 128, 425--467. averyetalRPranking
work page 2012
-
[4]
Review of Economic Studies, 81, 971--1002
Chade, Hector, Gregory Lewis, and Lones Smith (2014), Student Portfolios and the College Admissions Problem . Review of Economic Studies, 81, 971--1002. chade2014student
work page 2014
-
[5]
Chade, Hector and Lones Smith (2006), Simultaneous Search . Econometrica, 74, 1293--1307. chade2006simultaneous
work page 2006
-
[6]
Journal of Marketing Research, 56, 691--707
Dearden, James A, Rajdeep Grewal, and Gary L Lilien (2019), Strategic Manipulation of University Rankings, the Prestige Effect, and Student University Choice . Journal of Marketing Research, 56, 691--707. dearden2019strategic
work page 2019
-
[7]
Ehrenberg, Ronald G (2003), Method or Madness? Inside the "USNWR" College Rankings. ERIC. ehrenberg2003method
work page 2003
-
[8]
Journal of Political Economy, 122, 225--281
Fu, Chao (2014), Equilibrium Tuition, Applications, Admissions, and Enrollment in the College Market . Journal of Political Economy, 122, 225--281. fu2014equilibrium
work page 2014
Show all 19 references
-
[9]
American Economic Journal: Economic Policy, 12, 115--158
Goodman, Joshua, Oded Gurantz, and Jonathan Smith (2020), Take Two! SAT Retaking and College Enrollment Gaps . American Economic Journal: Economic Policy, 12, 115--158. goodman2020take
2020
-
[10]
Review of Economics and Statistics, 98, 671--684
Goodman, Sarena (2016), Learning from the Test: Raising Selective College Enrollment by Providing Information . Review of Economics and Statistics, 98, 671--684. goodman2016learning
2016
-
[11]
News Ranked Columbia No
Hartocollis, Anemona (2022), U.S. News Ranked Columbia No. 2, but a Math Professor Has His Doubts . The New York Times, 1, ://www.nytimes.com/2022/03/17/us/columbia-university-rank.html. columbia
2022
-
[12]
American Economic Journal: Applied Economics, 9, 223--61
MacLeod, W Bentley, Evan Riehl, Juan E Saavedra, and Miguel Urquiola (2017), The Big Sort: College Reputation and Labor Market Outcomes . American Economic Journal: Applied Economics, 9, 223--61. macleod2017big
2017
-
[13]
American Economic Review, 105, 3471--88
MacLeod, W Bentley and Miguel Urquiola (2015), Reputation and School Competition . American Economic Review, 105, 3471--88. macleod2015reputation
2015
-
[14]
Econometrica, 157--180
Milgrom, Paul and Chris Shannon (1994), Monotone Comparative Statics . Econometrica, 157--180. milgrom1994monotone
1994
-
[15]
NBER Working Paper, 2020--08
Mountjoy, Jack and Brent Hickman (2021), The Returns to College(s): Relative Value-Added and Match Effects in Higher Education . NBER Working Paper, 2020--08. mountjoy2021returns
2021
-
[16]
Whose Fault Is That? The New York Times, 5, ://www.nytimes.com/2023/09/05/magazine/college-worth-price.html
Tough, Paul (2023), Americans Are Losing Faith in the Value of College. Whose Fault Is That? The New York Times, 5, ://www.nytimes.com/2023/09/05/magazine/college-worth-price.html. tough2023americans
2023
-
[17]
://www.justice.gov/usao-edpa/pr/former-temple-business-school-dean-indicted-fraud
United States Attorney's Office (2021), Former Temple Business School Dean Indicted for Fraud . ://www.justice.gov/usao-edpa/pr/former-temple-business-school-dean-indicted-fraud. temple
2021
-
[18]
, " * write output.state after.block = add.period write newline
ENTRY address annote author booktitle chapter edition editor eid howpublished institution journal key month note number organization pages publisher school series title type url volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence a...
-
[19]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.