REVIEW 5 major objections 4 minor 14 references
Four Bottomless Errors and the Collapse of Statistical Fairness
T0 review · 5 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper argues that statistical fairness in AI is a category error, not a calibration problem, and that the field's accumulated work should be deleted.
desk verdict A sharp, readable provocation that earns attention but not its radical conclusion: the four fairness critiques are real, while the collapse thesis is stipulated rather than demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the "bottomless error": a fairness conception structured so that every attempted correction repeats the original mistake, in the paper's phrase, like answering 2 + 2 = apple rather than 2 + 2 = 5. The specific engine is the inversion "math creates ethics," where statistics and equality signs generate the ethical principle rather than an ethical principle governing the statistics. The paper uses this inversion to convert four contingent criticisms—equality, perspective, disproportion, and group fairness—into necessary ones, so that the impossibility of incremental repair follows from the nature of the errors rather than from any particular metric.
What would settle it
Find one published fairness metric whose recommendations were checked against an ethical principle stated independently of the metric, in a concrete high-stakes case; if the metric matches the independently stated principle, the claim that statistical fairness necessarily inverts ethics and math is falsified.
Extended reading notes
Core claim
The central claim is that the AI ethics of statistical fairness fails not in its implementations but in its starting point: it treats "What is fairness?" as a mathematical question rather than an ethical one. The four errors are presented as necessary consequences of that inversion. Conflating fairness with equality forbids discriminating judgments altogether; perspectival fairness defines one group's view by negating others; disproportion claims leverage suppressed base rates to generate misinformation; group fairness evaluates predefined social categories instead of letting fair treatment define groups. Since each error is constitutive of how statistical fairness works, resolving one difficulty only deepens it, so the approach collapses from within. The paper does not argue for better metrics; it argues that the project should be deleted, and that future work should begin from the principle that fairness treats equals equally and unequals proportionately unequally, with statistics applied afterward.
Load-bearing premise
The whole collapse rests on the premise that fairness's meaning must be fixed before any statistics are applied; if a statistical definition can legitimately inform what fairness means, the four errors become ordinary design problems instead of bottomless ones.
Editorial extensions
If this is right
- The COMPAS/ProPublica dispute is not a trade-off between competing fairness definitions; both sides are said to suppress one side of a numerical disproportion and thereby deny the ethical existence of those on the other side.
- Every common metric—statistical parity, predictive parity, false-positive error-rate balance, equalized odds, overall accuracy equality—is an instance of the equality error, so no combination of them can produce fairness.
- Fairness and accuracy are separable: an algorithm that is always wrong can still be perfectly fair if its errors are consistent, because fairness is about consistency, not accuracy.
- Fairness should produce groups rather than evaluate predefined ones, and the resulting groups may be counterintuitive or legally unrecognized; the paper suggests such surprises are a sign that ethical progress is happening.
Reading between the lines
- A testable extension not pursued in the paper: ask affected people to judge fairness before showing them any statistical definition, then compare their judgments with common fairness metrics; systematic disagreement would support the inversion claim, while agreement would undercut it.
- A formalization the author leaves implicit: the individual-consistency remedy can be stated as a pairwise condition—any two similar individuals should receive similar decisions—with group imbalances becoming descriptive outputs rather than inputs; this would inherit a new version of the equality error if "similar" is fixed by features rather than by fair process.
- If the collapse thesis is correct, the practical upshot is that fairness benchmarking is testing artifacts of the same category mistake, and auditing effort should move from group-parity constraints toward procedural review and individual consistency.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that the field of statistical fairness in AI ethics rests on four 'bottomless errors' (equality, perspective, disproportion, and group fairness), that these errors are integral to how statistical fairness works, and that consequently the approach cannot be corrected and should be abandoned, with the accumulated literature deleted. The argument is framed through a contrast between a method in which 'math creates ethics' and an Aristotelian conception in which the meaning of fairness is fixed before mathematics is applied. The paper also sketches alternative directions, including individual-level consistency and fairness as a process that creates groups.
Significance. If the central claim were established, the paper would have profound implications for a large and active research community. The paper is genuinely engaged with real sources, including Mehrabi et al., Verma and Rubin, Narayanan, Angwin et al., and Kleinberg, and it identifies some legitimate problems, such as the problematic equality-based definition in Mehrabi et al. and ProPublica's selective presentation of COMPAS statistics. However, the load-bearing inference from these examples to the collapse of the entire field is not supported. The evidence shows contingent defects in particular definitions, presentations, or journalistic choices, not errors that are integral to statistical fairness as such. Moreover, the foundational dichotomy between ethics-first and math-first reasoning is stipulated rather than defended. The significance of the paper as a rigorous argument is therefore limited, although it may serve as a provocative philosophical essay.
major comments (5)
- [Sections 2 and 3] The central dichotomy is stipulated rather than argued. Section 2 defines statistical fairness as a method in which 'math creates ethics,' and Section 3 concludes that 'Aristotle is ethics and statistical approaches are not.' The paper never refutes the possibility that statistical formalizations can legitimately inform, constrain, or partially constitute the meaning of fairness. Because this dichotomy is load-bearing for the claim that statistical fairness is 'not even on the continuum between right and wrong,' the collapse conclusion is substantially secured by stipulation.
- [Abstract and Sections 4–7] The claim that the four errors are 'integral to how statistical fairness works' is asserted but never demonstrated across the field. The evidence consists of selected examples: one definition from Mehrabi et al., Narayanan's 21 Fairness Definitions talk, ProPublica's COMPAS report, and Kleinberg's tradeoff presentation. Even if each example exhibited a genuine error, the inductive step from four examples to 'the larger project collapses from within' is unsupported. The paper does not show that these errors arise from the mathematical method itself rather than from particular choices made by particular authors.
- [Section 4] The claim that Verma and Rubin's five definitions all 'generate from an ethical logic moving in the wrong direction' because their names include parity, balance, or equality is inaccurate. Equalized odds and predictive parity are not equality-of-outcome constraints; they require equal error rates across groups, which can be a formalization of an antecedent moral claim that group membership should not change error rates. Mischaracterizing these definitions weakens the 'equality error' as a bottomless error and instead suggests a contestable design choice.
- [Section 6] The disproportion error is, by the paper's own account, an omission by ProPublica: 'the authors were correct about the numbers, but neglected to mention other numbers.' That is a defect in a particular piece of journalism or in a particular framing of results, not in statistical fairness as a method. A fixable reporting error provides no evidence that statistical fairness itself collapses from within, and the paper does not show that this error is reproduced necessarily by the method.
- [Section 7] The discussion of Kleinberg's tradeoff talk mischaracterizes its role. The tradeoff theorems show that several plausible fairness conditions cannot all hold simultaneously; that is a constraint on value choices, not evidence that group fairness as such is generated by statistics. The paper treats the existence of tradeoffs as proof that group fairness is a 'bottomless error,' but the theorems presuppose that the fairness criteria are ethically motivated, not that they are derived from mathematics.
minor comments (4)
- [Section 2] The word 'bazar' should be 'bizarre' in the sentence describing the moralized math example.
- [References] The manuscript contains several '[Redacted]' placeholders and the note 'References pending final inclusion/exclusion, formatting'; these must be resolved before any publication decision.
- [Figure 1] The caption states that the tables are 'conceptually correct, though the specific numbers are in some dispute (Barenstein 2019)'; this should be clarified with a more precise explanation of which numbers are disputed and how the figure should be interpreted.
- [Section 5] The paper's use of video timestamps is useful, but the corresponding reference entries should consistently include the date the video was accessed and, where available, a published transcript or slides.
Circularity Check
The collapse of statistical fairness is substantially secured by stipulation: Section 2 defines the method as 'math creates ethics' and Section 3 defines fairness as an a priori ethical concept, so the verdict that statistical fairness is not ethics follows by definition.
-
self definitional
[Section 3, paragraph 3, and Section 2, paragraph 2]
"What is significant is not what Aristotle’s definition means, instead, it is that the definition leads to debates about fairness within the parameters of ethics and reason. What this essay shows is that statistical approaches do not exist within those parameters. They have no relation – neither positive nor negative – with the ethics of fairness. So, it is not that Aristotle’s approach is better ethics than statistical fairness approaches, it is that Aristotle is ethics and statistical approaches are not."
The argument defines statistical fairness as 'a method [that] applies statistics to produce conceptions of fairness, as opposed to starting with a conception of fairness and then applying it statistically' (Section 2), and defines ethics/fairness as a concept fixed before mathematics, stating the need for 'grasping the meaning of fairness first, and by applying it algorithmically only subsequently' (Section 1). The conclusion that statistical fairness has 'no relation' with ethics is therefore contained in those definitions: the method's defining feature is the negation of the stipulated ethical procedure.
full rationale
The four error sections are not themselves circular: they critique specific definitions, talks, and a ProPublica analysis, and they include independent arguments (for example, equality as absence of favoritism prohibits any discriminating decision; the tradeoff talk shows conflicting group-level constraints). The redacted citations do not carry the collapse argument. However, the central inference from 'math creates ethics' to 'statistical fairness is not ethics' is self-definitional, because the paper stipulates both terms so that the conclusion follows by construction. The paper's assertion that the four errors are 'integral to how statistical fairness works' is also an unsupported generalization, but that is a correctness risk rather than a circularity. Overall, the collapse is substantially secured by stipulation, so a score of 6 is appropriate.
Assumptions & free parameters
free parameters (2)
- Selection of the four errors =
equality, perspective, disproportion, group
- Recidivism 'twice as likely' ratios =
2x (contested)
assumptions (5)
- domain assumption Aristotle's definition ("treat equals equally and unequals proportionately unequally") is the stable, correct account of fairness, and fairness is process-based and individual-first (Section 3).
- ad hoc to paper Ethical concepts must be fully determined before mathematics is applied; statistics-generated ethics is categorically wrong (Sections 2 and 3).
- domain assumption Narayanan's presentation is "fairness as constituted by discounting the experiences of others" (Section 5).
- domain assumption Calibration-style equality of decisions across groups is the correct standard for judging COMPAS, so ProPublica's finding is a "sleight-of-hand" (Section 6).
- domain assumption Fairness is about individuals before society (Section 3, via Nozick).
invented entities (1)
-
"Bottomless errors" as a category of wrongheaded error ("2 + 2 = apple")
Cite this review
Pith. "Pith review of Four Bottomless Errors and the Collapse of Statistical Fairness." pith.science (2026). https://pith.science/paper/XSH25JIM
@misc{pith2026250413790,
author = {Pith},
title = {Pith review of: Four Bottomless Errors and the Collapse of Statistical Fairness},
year = {2026},
howpublished = {\url{https://pith.science/paper/XSH25JIM}},
note = {Machine review of arXiv:2504.13790}
}
read the original abstract
The AI ethics of statistical fairness is an error, the approach should be abandoned, and the accumulated academic work deleted. The argument proceeds by identifying four recurring mistakes within statistical fairness. One conflates fairness with equality, which confines thinking to similars being treated similarly. The second and third errors derive from a perspectival ethical view which functions by negating others and their viewpoints. The final mistake constrains fairness to work within predefined social groups instead of allowing unconstrained fairness to subsequently define group composition. From the nature of these misconceptions, the larger argument follows. Because the errors are integral to how statistical fairness works, attempting to resolve the difficulties only deepens them. Consequently, the errors cannot be corrected without undermining the larger project, and statistical fairness collapses from within. While the collapse ends a failure in ethics, it also provokes distinct possibilities for fairness, data, and algorithms. Quickly indicating some of these directions is a secondary aim of the paper, and one that aligns with what fairness has consistently meant and done since Aristotle.
Figures
Reference graph
Works this paper leans on
-
[1]
Introduction Statistical fairness is today’s central approach to questions about the distribution of benefits and harms in AI ethics discussions when led by computer scientists – it is the way computer scientists do ethics (Carey and Wu 2023). It is also a particularly insidious Four Bottomless Errors and the Collapse of Statistical Fairness 2 kind of err...
work page 2023
-
[2]
What is statistical fairness? Statistical fairness can be defined superficially as a set of authors, articles, and conferences where the work appears. Many of the representative figures and documents are gathered in this paper, and together they form a movement, a common set of ideas and practices. More penetratingly, statistical fairness is a method. It ...
work page 2019
-
[3]
What is fairness? Fairness has been stably defined for nearly 2500 years: Treat equals equally and unequals proportionately unequally (Aristotle 350 BC: Book 5, 3A). While peripheral debates about the conception trace through philosophy’s history (Broome 1990), at least two central claims remain solid. First, fairness is about a process and not an outcome...
work page 1974
-
[4]
Equality error Fairness can be misconceived as straight equality – everyone treated the same – as in this example: Fairness is the absence of any prejudice or favoritism toward an individual or group based on their inherent or acquired characteristics. This particular rejection of prejudice and favoritism creates an insuperable problem: no discriminating ...
work page 2021
-
[5]
Perspective error 21 Fairness Definitions and Their Politics is a standard reference in statistical fairness (Castelnovo et al. 2022, Courtland 2018), and the presentation’s central use case is the Four Bottomless Errors and the Collapse of Statistical Fairness 6 COMPAS algorithmic crime predictor, which estimates whether an arrested defendant will reoffe...
work page 2022
-
[6]
Disproportion error Disproportion is a subset of the perspective error. Besides participating in the logic of exclusion as ethics, it positively leverages suppressed information in a dataset to generate misinformation. The paradigmatic example is ProPublica’s activist research publication dedicated to the bail or jail COMPAS algorithm. About it, ProPublic...
work page 2016
-
[7]
Group fairness error Because of the intersection with significant legal disputes and broad social justice, research in group fairness may be the most prominently debated area of statistical fairness (Mashiat et al. 2022, Chouldechova and Roth 2020, Binns 2020, Bird et al. 2020, Wachter et al. 2020). What is certain is that algorithms can be unfairly biase...
work page 2021
-
[11]
Propublica's compas data revisited
Track on Datasets and Benchmarks. Accessed 2 October 2024, https://datasets- benchmarks-proceedings.neurips.cc/paper/2021/file/92cc227532d17e56e07902b254dfad10- Paper-round1.pdf Barenstein, Matias. "Propublica's compas data revisited." arXiv preprint arXiv:1906.04711 (2019). Baumann, Joachim; Hannák, Anikó; and Christoph Heitz. 2022. Enforcing Group Fairn...
arXiv 2019
Show all 14 references
-
[12]
A snapshot of the frontiers of fairness in machine learning
Breast cancer. Accessed 1 November 2023, https://www.cdc.gov/cancer/breast/men/index.htm. Chouldechova, Alexandra, and Aaron Roth. 2020. "A snapshot of the frontiers of fairness in machine learning." Communications of the ACM 63, no. 5 (2020): 82-89. Cole, Matthew; Callum, Can...
2020
-
[13]
Trade-offs between group fairness metrics in societal resource allocation
Predictability and Surprise in Large Generative Models. Association for Computing Machinery, Conference on Fairness, Accountability, and Transparency (FAccT '22). DOI: https://doi.org/10.1145/3531146.3533229 Hao, Karen and Jonathan Stray. 2019. Can you make AI fairer than a ju...
-
[14]
Discrimination, Bias, Fairness, and Trustworthy AI
A Survey on Bias and Fairness in Machine Learning. ACM Comput. Surv. 54, 6, Article 115 (July 2022), 35 pages. DOI: https://doi.org/10.1145/3457607 Four Bottomless Errors and the Collapse of Statistical Fairness 15 Meng, Xiao-Li and Liberty Vittert (Hosts). (2022). Data Scienc...
2022
-
[23]
Bao, Michelle and Zhou, Angela and Zottola, Samantha and Brubach, Brian and Brubach, Brian and Desmarais, Sarah and Horowitz, Aaron and Lum, Kristian and Venkatasubramanian, Suresh
Accessed 1 February 2023, www.propublica.org/article/machine-biasrisk-assessments-in- criminal-sentencing. Bao, Michelle and Zhou, Angela and Zottola, Samantha and Brubach, Brian and Brubach, Brian and Desmarais, Sarah and Horowitz, Aaron and Lum, Kristian and Venkatasubramani...
2023
-
[2021]
35th Conference on Neural Information Processing Systems (NeurIPS
It’s COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks. 35th Conference on Neural Information Processing Systems (NeurIPS
-
[2022]
Parity” “balance
to medical researchers employing AI (Benjamin et al. 2024) to private industry (Varona and Suárez. 2022). So, the subject here is not a flawed detail infiltrating a few marginal publications. It is mainstream thought impacting the academics of computer science and the practica...
2024
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.