REVIEW 2 major objections 1 minor
Proxy Discrimination After Students for Fair Admissions
T0 review · 2 major / 1 minor · reviewed 2026-05-23 · grok-4.3
Pith's one-line read Decision tools are narrowly tailored when they exhibit the weakest total proxy power.
desk verdict The paper proposes a comparative 'weakest total proxy power' test for narrow tailoring in algorithmic proxies post-SFFA, but supplies no workable way to measure or compare that quantity. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Total proxy power, the aggregate measure of how strongly variables proxy for protected classes, which enables direct comparison of algorithms with matching accuracy rates.
What would settle it
An experiment showing that two algorithms with identical accuracy cannot be objectively ranked by total proxy power without introducing new forms of bias or subjective choices.
Extended reading notes
Core claim
Decision tools that use proxies are narrowly tailored when they exhibit the weakest total proxy power. The test is necessarily comparative. Thus, if two algorithms predict loan repayment or university academic performance with identical accuracy rates, but one uses zip code and the other does not, then the second algorithm can be said to have deployed a more equitable means for achieving the same result as the first algorithm. Lawmakers can develop caps to permissible proxy power over time as courts and algorithm builders learn more about the power of variables.
Load-bearing premise
That total proxy power is a measurable and objective quantity that can be consistently compared across algorithms without post-hoc manipulation or new forms of bias.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that post-Students for Fair Admissions there is no clear legal test for regulating proxy variables for protected classes in decision tools. It proposes that such tools are narrowly tailored when they exhibit the weakest total proxy power; the test is necessarily comparative, so that equal-accuracy algorithms can be ranked by whether one uses a proxy such as zip code. The paper suggests lawmakers can develop caps on permissible proxy power over time and argues that plaintiffs should bear the burden of producing less discriminatory alternatives provided testing data is available.
Significance. If operationalized, the comparative test could supply courts and regulators with a workable standard for evaluating proxy use in algorithmic lending, admissions, and similar domains, potentially reducing reliance on ad-hoc disparate-impact analyses while preserving predictive accuracy.
major comments (2)
- [Abstract] Abstract and the section developing the test: the central claim that algorithms can be ranked by 'weakest total proxy power' when accuracy is held constant requires an operational definition, loss function, or aggregation rule for computing this scalar from data. No such definition, measurement procedure, or edge-case analysis is supplied, rendering the comparative test non-applicable without embedding contested normative choices about which correlations count as proxy power.
- [Discussion of caps] The paragraph on developing caps over time: the suggestion that courts and builders can learn permissible levels of proxy power assumes an objective, comparable metric exists and can be refined empirically, yet the manuscript provides neither a baseline measurement protocol nor a falsifiable procedure for updating caps.
minor comments (1)
- [Abstract] The abstract states the test is 'necessarily comparative' but does not clarify how non-identical but comparable results are to be handled; a brief illustrative example or decision tree would improve clarity.
Simulated Author's Rebuttal
We thank the referee for these constructive comments on operationalizing the proposed test. We respond to each major comment below and indicate where revisions will be made to clarify the legal nature of the framework.
read point-by-point responses
-
Referee: [Abstract] Abstract and the section developing the test: the central claim that algorithms can be ranked by 'weakest total proxy power' when accuracy is held constant requires an operational definition, loss function, or aggregation rule for computing this scalar from data. No such definition, measurement procedure, or edge-case analysis is supplied, rendering the comparative test non-applicable without embedding contested normative choices about which correlations count as proxy power.
Authors: The test is a legal standard for narrow tailoring, not a computational metric. 'Weakest total proxy power' denotes a comparative judicial inquiry: when two algorithms achieve equal accuracy, the one that avoids variables functioning as proxies for protected classes (identified via existing statistical evidence and case law on disparate impact) is less discriminatory. No single loss function or scalar is supplied because the framework relies on litigants presenting evidence of proxy relationships, as courts already do in equal protection and Title VII cases. We will revise the abstract and test-development section to explicitly distinguish this legal comparative approach from a technical aggregation rule and to note that normative choices about proxy identification remain with the factfinder. revision: partial
-
Referee: [Discussion of caps] The paragraph on developing caps over time: the suggestion that courts and builders can learn permissible levels of proxy power assumes an objective, comparable metric exists and can be refined empirically, yet the manuscript provides neither a baseline measurement protocol nor a falsifiable procedure for updating caps.
Authors: The discussion of caps is forward-looking and does not assert that an objective metric currently exists. It posits that repeated application of the comparative test across cases can generate data from which regulators or legislatures might derive permissible thresholds, analogous to how other legal standards have evolved. We agree the manuscript would benefit from greater caution here and will revise the paragraph to clarify that any cap development would require external empirical work and to acknowledge the absence of a specific measurement protocol in the current proposal. revision: yes
Circularity Check
No circularity; proposal is a direct normative definition without self-referential reduction
full rationale
The paper states its central test directly as a legal standard ('Decision tools that use proxies are narrowly tailored when they exhibit the weakest total proxy power') without any equations, fitted parameters, or derivations that reduce the claim to its own inputs by construction. No self-citations are invoked as load-bearing uniqueness theorems, no ansatzes are smuggled, and the comparative framing is presented as a definitional choice rather than an output forced by prior results. The absence of an operational metric for 'total proxy power' is a completeness issue, not a circularity issue, as the text does not claim to derive the test from independent premises that loop back to the same definition.
Assumptions & free parameters
assumptions (1)
- domain assumption Supreme Court precedent in Students for Fair Admissions requires narrow tailoring for any race-conscious or proxy-based policies.
Cite this review
Pith. "Pith review of Proxy Discrimination After Students for Fair Admissions." pith.science (2026). https://pith.science/paper/2501.03946
@misc{pith2026250103946,
author = {Pith},
title = {Pith review of: Proxy Discrimination After Students for Fair Admissions},
year = {2026},
howpublished = {\url{https://pith.science/paper/2501.03946}},
note = {Machine review of arXiv:2501.03946}
}
read the original abstract
Today, there is no clear legal test for regulating the use of variables that proxy for race and other protected classes and classifications. This Article develops such a test. Decision tools that use proxies are narrowly tailored when they exhibit the weakest total proxy power. The test is necessarily comparative. Thus, if two algorithms predict loan repayment or university academic performance with identical accuracy rates, but one uses zip code and the other does not, then the second algorithm can be said to have deployed a more equitable means for achieving the same result as the first algorithm. Scenarios in which two algorithms produce comparable and non-identical results present a greater challenge. This Article suggests that lawmakers can develop caps to permissible proxy power over time, as courts and algorithm builders learn more about the power of variables. Finally, the Article considers who should bear the burden of producing less discriminatory alternatives and suggests plaintiffs remain in the best position to keep defendants honest - so long as testing data is made available.
Reviewed May 23, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.