REVIEW 4 major objections 2 minor 137 references
The Fair Game: Auditing & Debiasing AI Algorithms Over Time
T0 review · 4 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper claims that fairness in machine learning can be kept adaptive by an RL loop where an auditor's changing bias definition automatically retunes a debiasing algorithm, without model redesign.
desk verdict The abstract proposes a genuinely interesting adjustable-auditor idea for fair ML, but the full text is an unrelated superconductivity paper, so there is nothing to review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Auditor–Debiaser loop under reinforcement learning: the Auditor's bias measure serves as the reward signal for the RL debiaser, which alters the ML model's predictions. Because the reward is redefined by the Auditor, the fairness objective is non-stationary and can be swapped without touching the underlying model.
What would settle it
Simulate the loop with an auditor that alternates between two conflicting definitions, such as demographic parity and equalized odds, and measure the model's error and fairness against ground-truth labels; if the model tracks the auditor but ground-truth fairness degrades or the loop oscillates, the central claim of reliable adaptation fails.
Extended reading notes
Core claim
Fair Game is a two-component RL loop wrapped around any ML algorithm. An Auditor quantifies bias using a chosen observational definition and emits feedback; a Debiasing algorithm, driven by reinforcement learning, adjusts the ML predictions to minimize that defined unfairness. The author's key claim is that to change the fairness criterion, one only replaces or modifies the Auditor—the feedback it sends changes, and the RL debiaser adapts accordingly. Thus the same loop works pre- and post-deployment, and can track changing societal or legal fairness norms over time.
Load-bearing premise
The loop can only be fair if the auditor's observational bias measure is a reliable and sufficient signal for steering the model toward genuine fairness, yet the abstract itself notes such measures need ground truth or hindsight; the loop's adaptivity does not resolve that grounding gap.
Editorial extensions
If this is right
- Deployed ML models can be kept fair as legal or ethical fairness criteria change, without model redesign.
- Conflicting bias definitions can be handled by switching the auditor's metric rather than retraining the model.
- Fairness is maintained both before and after deployment, addressing the retrospective limitation of observational measures.
- The RL debiaser can track a non-stationary reward target, enabling simulation of evolving social norms.
Reading between the lines
- The meaningfulness of the loop depends on the auditor's metric being a valid proxy for the fairness one actually cares about; an arbitrary auditor would define fairness circularly.
- If the auditor's metric changes abruptly, the RL loop may become unstable; the paper does not address convergence under time-varying rewards.
- The supplied full text is a physics manuscript on nickelate superconductors, not the described Fair Game framework, so the abstract's claims are not supported by derivations in the provided document.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submitted manuscript, arXiv:2508.06443, titled 'The Fair Game: Auditing & Debiasing AI Algorithms Over Time,' proposes a framework in which an Auditor and a Debiasing algorithm are placed in a reinforcement learning loop around an ML algorithm. The abstract claims that this loop can 'assure fairness' and adapt predictions over time by modifying only the auditor, thereby simulating evolving ethical and legal frameworks. However, the provided full text is a condensed-matter paper on the multiorbital character of density waves in trilayer nickelate superconductors. It contains no mention of Fair Game, Auditor, Debiasing, reinforcement learning, fairness, or any component of the claimed framework. Thus the central claims are unsupported by any technical specification, derivation, or experimental validation.
Significance. If the proposed framework were operational, dynamic adaptation of fairness criteria via an auditor-debiaser RL loop would be a potentially significant step in fair machine learning. The abstract identifies a real limitation of observational bias definitions and proposes a plausible direction. However, the submitted manuscript provides no equations, no algorithm definition, no stability or convergence analysis, and no experiments. Because the full text is an unrelated superconductivity paper, the claimed contribution is entirely unverifiable. The paper therefore cannot be assessed as a scientific contribution in its current form.
major comments (4)
- [Full text] The full text of the manuscript is an unrelated condensed-matter paper on trilayer nickelate superconductors. It contains no discussion of Fair Game, the Auditor, the Debiasing algorithm, reinforcement learning, fairness metrics, or any of the components named in the abstract. This is a load-bearing defect: the central claim of an adaptive fairness loop has no supporting specification whatsoever. Neither a reader nor a referee can verify the proposed mechanism, and there is no basis for judging correctness, novelty, or reproducibility.
- [Abstract (observational bias limitations)] The abstract states that observational bias definitions 'can only be deployed if either the ground truth is known or only in retrospect,' yet the proposed loop feeds exactly such observational measures into the debiasing agent. The abstract does not explain how the auditor obtains ground truth, how it avoids the retrospective gap, or why optimizing against the auditor's own criteria constitutes genuine fairness rather than circular self-confirmation. This is not merely a presentation issue; it undermines the central claim that the framework can 'assure fairness' rather than internally defined compliance.
- [No algorithm specification or analysis] No equations, pseudocode, or formal description of the Fair Game loop are given. There is no specification of the state space, action space, reward signal, policy optimization method, convergence conditions, or stability under a changing reward target. Without these, the core claim that RL can adapt fairness goals over time 'by only modifying the auditor' cannot be evaluated, and the non-stationary reward problem introduced by a time-varying auditor is unaddressed.
- [No experiments or comparisons] The manuscript contains no empirical evaluation, no case studies, and no comparison to existing fair-ML or RL-for-fairness methods. The abstract's claims about pre- and post-deployment adaptation are therefore unsupported. This would be a major shortcoming even if the full text were on-topic, and it is decisive given the absence of any technical content.
minor comments (2)
- [Abstract] The abstract contains awkward repetitions, e.g., the name 'Fair Game' appearing twice in overlapping sentences, and undefined terms such as 'Auditor' and 'Debiasing algorithm' that are not described anywhere in the supplied text. These presentation issues are secondary to the substantive defects.
- [General] If the full-text mismatch is due to an administrative error, the authors should resubmit the correct manuscript with complete technical appendices. As it stands, the document is internally inconsistent: title and abstract do not correspond to the body.
Circularity Check
No circular derivation present: the full text is an unrelated superconductivity paper, and the abstract's Fair Game framework is asserted rather than derived.
full rationale
The provided full text is a condensed-matter paper on trilayer nickelate superconductors; it does not contain the Auditor, Debiasing algorithm, reinforcement-learning loop, or any equations or results belonging to the 'Fair Game' framework. There is therefore no derivation chain in which a claimed result can be shown to reduce to its inputs. The abstract's concession that observational bias definitions require ground truth or retrospective knowledge identifies a real limitation, and the proposed loop does not resolve that limitation, but that is a support/correctness gap rather than a circularity: the paper does not derive 'fairness' from the auditor's measure by equation or by construction. The statement that fairness goals can be adapted 'by only modifying the auditor' describes a feedback architecture, not a logical reduction of a conclusion to its premise. Without any actual derivation, no specific circular step can be exhibited, so the honest finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- Auditor's bias criteria and their thresholds
assumptions (3)
- domain assumption The RL debiasing loop learns a stable policy while the auditor's reward-relevant criteria change over time
- domain assumption Observational bias measures quantified by the auditor are sufficient to drive genuine fairness improvement
- domain assumption Modifying the auditor over time simulates the evolution of ethical and legal frameworks
invented entities (2)
-
The Auditor
-
The Debiasing algorithm (RL agent)
Cite this review
Pith. "Pith review of The Fair Game: Auditing & Debiasing AI Algorithms Over Time." pith.science (2026). https://pith.science/paper/VTFX5AKO
@misc{pith2026250806443,
author = {Pith},
title = {Pith review of: The Fair Game: Auditing & Debiasing AI Algorithms Over Time},
year = {2026},
howpublished = {\url{https://pith.science/paper/VTFX5AKO}},
note = {Machine review of arXiv:2508.06443}
}
read the original abstract
An emerging field of AI, namely Fair Machine Learning (ML), aims to quantify different types of bias (also known as unfairness) exhibited in the predictions of ML algorithms, and to design new algorithms to mitigate them. Often, the definitions of bias used in the literature are observational, i.e. they use the input and output of a pre-trained algorithm to quantify a bias under concern. In reality,these definitions are often conflicting in nature and can only be deployed if either the ground truth is known or only in retrospect after deploying the algorithm. Thus,there is a gap between what we want Fair ML to achieve and what it does in a dynamic social environment. Hence, we propose an alternative dynamic mechanism,"Fair Game",to assure fairness in the predictions of an ML algorithm and to adapt its predictions as the society interacts with the algorithm over time. "Fair Game" puts together an Auditor and a Debiasing algorithm in a loop around an ML algorithm. The "Fair Game" puts these two components in a loop by leveraging Reinforcement Learning (RL). RL algorithms interact with an environment to take decisions, which yields new observations (also known as data/feedback) from the environment and in turn, adapts future decisions. RL is already used in algorithms with pre-fixed long-term fairness goals. "Fair Game" provides a unique framework where the fairness goals can be adapted over time by only modifying the auditor and the different biases it quantifies. Thus,"Fair Game" aims to simulate the evolution of ethical and legal frameworks in the society by creating an auditor which sends feedback to a debiasing algorithm deployed around an ML system. This allows us to develop a flexible and adaptive-over-time framework to build Fair ML systems pre- and post-deployment.
Reference graph
Works this paper leans on
-
[1]
Agarwal, A., Beygelzimer, A., Dud \' k, M., Langford, J., and Wallach, H. (2018). A reductions approach to fair classification. In International conference on machine learning , pages 60--69. PMLR
2018
-
[2]
Ajarra, A., Ghosh, B., and Basu, D. (2024). Active fourier auditor for estimating distributional properties of ml models. arXiv preprint arXiv:2410.08111
work page Pith review arXiv 2024
-
[3]
Albarghouthi, A., D'Antoni, L., Drews, S., and Nori, A. V. (2017). Fairsquare: probabilistic verification of program fairness. Proc. ACM Program. Lang. , 1(OOPSLA)
2017
-
[4]
Altman, E. (1999). Constrained Markov decision processes https://hal.inria.fr/docs/00/07/41/09/PS/RR-2574.ps , volume 7. CRC Press
1999
-
[5]
(23 May 2016)
Angwin, J., Larson, J., Mattu, S., and Kirchner, L. (23 May 2016). Machine bias: There’s software used across the country to predict future criminals. And it’s biased against blacks, http://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing ProPublica
2016
-
[6]
Annas, G. J. (2003). Hipaa regulations: a new era of medical-record privacy? New England Journal of Medicine , 348:1486
2003
-
[7]
Arrow, K. (1971). The theory of discrimination
1971
-
[8]
Assaf, D., Gutman, Y., Neuman, Y., Segal, G., Amit, S., Gefen-Halevi, S., Shilo, N., Epstein, A., Mor-Cohen, R., Biber, A., et al. (2020). Utilization of machine-learning models to accurately predict the risk for critical covid-19. Internal and emergency medicine , 15:1435--1443
2020
Show all 137 references
-
[9]
and Basu, D
Azize, A. and Basu, D. (2022). When Privacy Meets Partial Information: A Refined Analysis of Differentially Private Bandits . In Advances in Neural Information Processing Systems , New Orleans, United States
2022
-
[10]
and Basu, D
Azize, A. and Basu, D. (2024). Concentrated Differential Privacy for Bandits . In 2024 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages 78--109, Toronto, Canada. IEEE , IEEE
2024
-
[11]
A., and Basu, D
Azize, A., Jourdan, M., Marjani, A. A., and Basu, D. (2023). On the Complexity of Differentially Private Best-Arm Identification with Fixed Confidence . In NeurIPS 2023 -- Conference on Neural Information Processing Systems , volume 36, pages 71150--71194, New Orleans (US), Un...
2023
-
[12]
Bagaric, M., Hunter, D., and Stobbs, N. (2019). Erasing the bias against using artificial intelligence to predict future criminality: algorithms are color blind and never tire. U. Cin. L. Rev. , 88:1037
2019
-
[13]
Bai, Y., Jin, C., Mei, S., and Yu, T. (2022). Near-optimal learning of extensive-form games with imperfect information. In International Conference on Machine Learning , pages 1337--1382. PMLR
2022
-
[14]
Barocas, S., Hardt, M., and Narayanan, A. (2023). Fairness and machine learning: Limitations and opportunities . MIT Press
2023
-
[15]
Bastani, O., Zhang, X., and Solar-Lezama, A. (2019). Probabilistic verification of fairness properties via concentration. Proceedings of the ACM on Programming Languages , 3(OOPSLA):1--27
2019
-
[16]
Basu, D., Dimitrakakis, C., and Tossou, A. (2020). Privacy in Multi-armed Bandits: Fundamental Definitions and Lower Bounds https://arxiv.org/pdf/1905.12298.pdf. NeurIPS Workshop on Privacy Preserving Machine Learning
2020 arXiv
-
[17]
Basu, D., Maillard, O.-A., and Mathieu, T. (2022). Bandits corrupted by Nature: Lower bounds on regret and robust optimistic algorithm https://arxiv.org/pdf/2203.03186.pdf. arXiv preprint arXiv:2203.03186 . (Submitted to COLT'22)
2022 arXiv
-
[18]
K., Dey, K., Hind, M., Hoffman, S
Bellamy, R. K., Dey, K., Hind, M., Hoffman, S. C., Houde, S., Kannan, K., Lohia, P., Martino, J., Mehta, S., Mojsilovi \'c , A., et al. (2019). Ai fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias. IBM Journal of Research and Development , 63(4/5):4--1
2019
-
[19]
Bellamy, R. K. E., Dey, K., Hind, M., Hoffman, S. C., Houde, S., Kannan, K., Lohia, P., Martino, J., Mehta, S., Mojsilovic, A., Nagar, S., Ramamurthy, K. N., Richards, J., Saha, D., Sattigeri, P., Singh, M., Varshney, K. R., and Zhang, Y. (2018). Ai fairness 360: An extensible...
2018
-
[20]
Blumrosen, A. W. (1967). The duty of fair recruitment under the civil rights act of 1964. Rutgers L. Rev. , 22:465
1967
-
[21]
Borca-Tasciuc, G., Guo, X., Bak, S., and Skiena, S. (2022). Provable fairness for neural network models using formal verification
2022
-
[22]
Brown, G. W. (1951). Iterative solution of games by fictitious play. Act. Anal. Prod Allocation , 13(1):374
1951
-
[23]
K., Segal, M., Basu, D., George, A.-M., and Dimitrakakis, C
Buening, T. K., Segal, M., Basu, D., George, A.-M., and Dimitrakakis, C. (2022). On Meritocracy in Optimal Set Selection . In EAAMO 2022- Equity and Access in Algorithms, Mechanisms, and Optimization , Arlington, United States. ACM
2022
-
[24]
Bénesse, C., Gamboa, F., Loubes, J.-M., and Boissin, T. (2021). Fairness seen as global sensitivity analysis
2021
-
[25]
and Z liobait \.e , I
Calders, T. and Z liobait \.e , I. (2013). Why unbiased computational processes can lead to discriminative decision procedures. In Discrimination and Privacy in the Information Society: Data mining and profiling in large databases , pages 43--57. Springer
2013
-
[26]
D., and Dubhashi, D
Carlsson, E., Basu, D., Johansson, F. D., and Dubhashi, D. (2024). Pure Exploration in Bandits with Linear Constraints . In International Conference on Artificial Intelligence and Statistics , volume 238 of Proceedings of Machine Learning Research (PMLR) , pages 334--342, Vale...
2024
-
[27]
cas, S., Hardt, M., and Narayanan, A. (2023). Fairness and Machine Learning: Limitations and Opportunities . MIT Press
2023
-
[28]
and Haas, C
Caton, S. and Haas, C. (2024). Fairness in machine learning: A survey. ACM Computing Surveys , 56(7):1--38
2024
-
[29]
E., Huang, L., Keswani, V., and Vishnoi, N
Celis, L. E., Huang, L., Keswani, V., and Vishnoi, N. K. (2019). Classification with fairness constraints: A meta-algorithm with provable guarantees. In Proceedings of the conference on fairness, accountability, and transparency , pages 319--328
2019
-
[30]
Chaudhuri, K., Monteleoni, C., and Sarwate, A. D. (2011). Differentially private empirical risk minimization. Journal of Machine Learning Research , 12(3)
2011
-
[31]
R., and Liu, H
Cheng, L., Varshney, K. R., and Liu, H. (2021). Socially responsible AI algorithms: Issues, purposes, and challenges. Journal of Artificial Intelligence Research , 71:1137--1181
2021
-
[32]
Cherian, J. J. and Cand \`e s, E. J. (2024). Statistical inference for fairness auditing. Journal of Machine Learning Research , 25(149):1--49
2024
-
[33]
Chierichetti, F., Kumar, R., Lattanzi, S., and Vassilvtiskii, S. (2019). Matroids, matchings, and fairness. In The 22nd international conference on artificial intelligence and statistics , pages 2212--2220. PMLR
2019
-
[34]
Chouldechova, A. (2017). Fair prediction with disparate impact: A study of bias in recidivism prediction instruments. Big data , 5(2):153--163
2017
-
[35]
and Roth, A
Chouldechova, A. and Roth, A. (2018). The frontiers of fairness in machine learning. arXiv preprint arXiv:1810.08810
2018 arXiv
-
[36]
and Roth, A
Chouldechova, A. and Roth, A. (2020). A snapshot of the frontiers of fairness in machine learning https://dl.acm.org/doi/pdf/10.1145/3376898. Communications of the ACM , 63(5):82--89
2020 doi
-
[37]
F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in neural information processing systems , 30
2017
-
[38]
Chzhen, E., Denis, C., Hebiri, M., Oneto, L., and Pontil, M. (2020). Fair regression with Wasserstein barycenters https://arxiv.org/pdf/2006.07286.pdf. Advances in Neural Information Processing Systems , 33:7321--7331
2020 arXiv
-
[39]
H., Jacobs, B
Conitzer, V., Freedman, R., Heitzig, J., Holliday, W. H., Jacobs, B. M., Lambert, N., Moss \'e , M., Pacuit, E., Russell, S., Schoelkopf, H., et al. (2024). Position: Social choice should guide ai alignment in dealing with diverse human feedback. In Forty-first International C...
2024
-
[40]
Corbett-Davies, S., Pierson, E., Feller, A., Goel, S., and Huq, A. (2017). Algorithmic decision making and the cost of fairness. In Proceedings of the 23rd acm sigkdd international conference on knowledge discovery and data mining , pages 797--806
2017
-
[41]
Cotter, A., Jiang, H., Gupta, M., Wang, S., Narayan, T., You, S., and Sridharan, K. (2019). Optimization with non-differentiable constraints with applications to fairness, recall, churn, and other goals. Journal of Machine Learning Research , 20(172):1--59
2019
-
[42]
D a browski, . D. and Suska, M. (2022). The European Union Digital Single Market: Europe's Digital Transformation . Routledge
2022
-
[43]
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y. (2023). Safe rlhf: Safe reinforcement learning from human feedback. arXiv preprint arXiv:2310.12773
2023 arXiv
-
[44]
Dandekar, A., Basu, D., and Bressan, S. (2018). Differential privacy for regularised linear regression. In International Conference on Database and Expert Systems Applications , pages 483--491. Springer
2018
-
[45]
and Basu, D
Das, U. and Basu, D. (2024). Learning to explore with lagrangians for bandits under unknown constraints. In Seventeenth European Workshop on Reinforcement Learning
2024
-
[46]
Daskalakis, C., Golowich, N., and Zhang, K. (2023). The complexity of markov equilibrium in stochastic games. In The Thirty Sixth Annual Conference on Learning Theory , pages 4180--4234. PMLR
2023
-
[47]
Devroye, L., Gy \"o rfi, L., and Lugosi, G. (2013). A probabilistic theory of pattern recognition , volume 31. Springer Science & Business Media
2013
-
[48]
and Roth, A
Dwork, C. and Roth, A. (2014). The algorithmic foundations of differential privacy https://www.tau.ac.il/ saharon/BigData2018/privacybook.pdf. Found. Trends Theor. Comput. Sci. , 9(3-4):211--407
2014
-
[49]
Elie, R., Perolat, J., Lauri \`e re, M., Geist, M., and Pietquin, O. (2020). On the convergence of model free learning in mean field games. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 34, pages 7143--7150
2020
-
[50]
Eriksson, H., Basu, D., Alibeigi, M., and Dimitrakakis, C. (2022a). Risk-Sensitive Bayesian Games for Multi-Agent Reinforcement Learning under Policy Uncertainty . In OptLearnMAS@AAMAS , Workshop on Optimization and Learning in Multiagent Systems at International Conference on...
-
[51]
Eriksson, H., Basu, D., Alibeigi, M., and Dimitrakakis, C. (2022b). SENTINEL: Taming Uncertainty with Ensemble-based Distributional Reinforcement Learning . In UAI 2022- Proceedings of the Thirty-Eighth Conference on Uncertainty in Artificial Intelligence , volume 180 of Proce...
2022
-
[52]
A., Moeller, J., Scheidegger, C., and Venkatasubramanian, S
Feldman, M., Friedler, S. A., Moeller, J., Scheidegger, C., and Venkatasubramanian, S. (2015). Certifying and removing disparate impact. In proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining , pages 259--268
2015
-
[53]
Feldman, V., Guruswami, V., Raghavendra, P., and Wu, Y. (2012). Agnostic learning of monomials by halfspaces is hard. SIAM Journal on Computing , 41(6):1558--1590
2012
-
[54]
Fiegel, C., M \'e nard, P., Kozuno, T., Munos, R., Perchet, V., and Valko, M. (2023). Adapting to game trees in zero-sum imperfect information games. In International Conference on Machine Learning , pages 10093--10135. PMLR
2023
-
[55]
Fiss, O. M. (1970). A theory of fair employment laws. U. Chi. L. Rev. , 38:235
1970
-
[56]
and Basu, D
Flet-Berliac, Y. and Basu, D. (2022). SAAC: Safe Reinforcement Learning as an Adversarial Game of Actor-Critics . In RLDM 2022 - The Multi-disciplinary Conference on Reinforcement Learning and Decision Making , Providence, United States. Accepted at the 5th Multi-disciplinary ...
2022
-
[57]
Gajane, P., Saxena, A., Tavakol, M., Fletcher, G., and Pechenizkiy, M. (2022). Survey on fair reinforcement learning: Theory and practice. arXiv preprint arXiv:2205.10032
2022 arXiv
-
[58]
Galhotra, S., Brun, Y., and Meliou, A. (2017). Fairness testing: testing software for discrimination. In Proceedings of the 2017 11th Joint meeting on foundations of software engineering , pages 498--510
2017
-
[59]
and Tamayo, P
Galindo, J. and Tamayo, P. (2000). Credit risk assessment using statistical and machine learning: basic methodology and risk modeling applications. Computational economics , 15:107--143
2000
-
[60]
Ghosh, B., Basu, D., and Meel, K. (2023a). ''How Biased are Your Features?'': Computing Fairness Influence Functions with Global Sensitivity Analysis . In FAccT '23: the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages 138--148, Chicago IL, United States. ACM
2023
-
[61]
Ghosh, B., Basu, D., and Meel, K. S. (2021). Justicia: A stochastic sat approach to formally verify fairness. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 35, pages 7554--7563
2021
-
[62]
Ghosh, B., Basu, D., and Meel, K. S. (2022a). Algorithmic fairness verification with graphical models. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 36, pages 9539--9548
-
[63]
Ghosh, B., Basu, D., and Meel, K. S. (2022b). Algorithmic fairness verification with graphical models. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 36, pages 9539--9548
-
[64]
Ghosh, B., Basu, D., and Meel, K. S. (2022c). Algorithmic fairness verification with graphical models . In AAAI-2022 - 36th AAAI Conference on Artificial Intelligence , volume 2, Virtual, United States
2022
-
[65]
Ghosh, B., Basu, D., and Meel, K. S. (2023b). ``how biased are your features?”: Computing fairness influence functions with global sensitivity analysis. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages 138--148
2023
-
[66]
Giannou, A., Lotidis, K., Mertikopoulos, P., and Vlatakis-Gkaragkounis, E.-V. (2022). On the convergence of policy gradient methods to nash equilibria in general stochastic games. Advances in Neural Information Processing Systems , 35:7128--7141
2022
-
[67]
Godinot, A., Le Merrer, E., Tr \'e dan, G., Penzo, C., and Ta \"i ani, F. (2023). Change-Relaxed Active Fairness Auditing . In RJCIA 2023 - 21e Rencontres des Jeunes Chercheurs en Intelligence Artificiel , CNIA, pages 91--96, Strasbourg, France. Association Fran c aise pour l'...
2023
-
[68]
Godinot, A., Le Merrer, E., Tr \'e dan, G., Penzo, C., and Ta \" ani, F. (2024). Under manipulations, are some ai models harder to audit? In 2024 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages 644--664. IEEE
2024
-
[69]
N., Shafer, J., and Yehudayoff, A
Goldwasser, S., Rothblum, G. N., Shafer, J., and Yehudayoff, A. (2021). Interactive proofs for verifying machine learning. In 12th Innovations in Theoretical Computer Science Conference (ITCS 2021) . Schloss-Dagstuhl-Leibniz Zentrum f \"u r Informatik
2021
-
[70]
Gordaliza, P., Del Barrio, E., Fabrice, G., and Loubes, J.-M. (2019). Obtaining fairness using optimal transport theory. In International conference on machine learning , pages 2357--2365. PMLR
2019
-
[71]
Gorti, A., Gaur, M., and Chadha, A. (2024). Unboxing occupational bias: Grounded debiasing llms with us labor data. arXiv preprint arXiv:2408.11247
2024 arXiv
-
[72]
Gu, S., Holly, E., Lillicrap, T., and Levine, S. (2017). Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In 2017 IEEE international conference on robotics and automation (ICRA) , pages 3389--3396. IEEE
2017
-
[73]
Gy \"o rfi, L., Kohler, M., Krzyzak, A., and Walk, H. (2006). A distribution-free theory of nonparametric regression . Springer Science & Business Media
2006
-
[74]
and Domingo-Ferrer, J
Hajian, S. and Domingo-Ferrer, J. (2012). A methodology for direct and indirect discrimination prevention in data mining. IEEE transactions on knowledge and data engineering , 25(7):1445--1459
2012
-
[75]
Hardt, M., Price, E., and Srebro, N. (2016). Equality of opportunity in supervised learning. Advances in neural information processing systems , 29
2016
-
[76]
H \'e bert-Johnson, U., Kim, M., Reingold, O., and Rothblum, G. (2018). Multicalibration: Calibration for the (computationally-identifiable) masses. In International Conference on Machine Learning , pages 1939--1948. PMLR
2018
-
[77]
and Krause, A
Heidari, H. and Krause, A. (2018). Preventing disparate treatment in sequential decision making. In IJCAI , pages 2248--2254
2018
-
[78]
M., Harman, M., and Sarro, F
Hort, M., Chen, Z., Zhang, J. M., Harman, M., and Sarro, F. (2024). Bias mitigation for machine learning classifiers: A comprehensive survey. ACM Journal on Responsible Computing , 1(2):1--52
2024
-
[79]
Huber, P. J. (1981). Robust statistics. Wiley Series in Probability and Mathematical Statistics
1981
-
[80]
Janaro, C. (2023). Nyc local law 144: A failed attempt at regulating ai in hiring
2023
-
[81]
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y. (2024). Beavertails: Towards improved safety alignment of llm via a human-preference dataset. Advances in Neural Information Processing Systems , 36
2024
-
[82]
and Nachum, O
Jiang, H. and Nachum, O. (2020). Identifying and correcting label bias in machine learning. In International conference on artificial intelligence and statistics , pages 702--712. PMLR
2020
-
[83]
and Calders, T
Kamiran, F. and Calders, T. (2012). Data preprocessing techniques for classification without discrimination. Knowledge and information systems , 33(1):1--33
2012
-
[84]
Kamishima, T., Akaho, S., Asoh, H., and Sakuma, J. (2012). Fairness-aware classifier with prejudice remover regularizer. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2012, Bristol, UK, September 24-28, 2012. Proceedings, Part II 23 ,...
2012
-
[85]
Kidder, W. C. (2001). Does the lsat mirror or magnify racial and ethnic differences in educational attainment: A study of equally achieving elite college students. Calif. L. Rev. , 89:1055
2001
-
[86]
Kim, M., Reingold, O., and Rothblum, G. (2018). Fairness through computationally-bounded awareness. Advances in neural information processing systems , 31
2018
-
[87]
Kleinberg, J., Mullainathan, S., and Raghavan, M. (2016). Inherent trade-offs in the fair determination of risk scores. arXiv preprint arXiv:1609.05807
2016 arXiv
-
[88]
Kleine Buening, T., Segal, M., Basu, D., George, A.-M., and Dimitrakakis, C. (2022). On meritocracy in optimal set selection. In Equity and Access in Algorithms, Mechanisms, and Optimization , pages 1--14. ACM
2022
-
[89]
and Szepesv \'a ri, C
Lattimore, T. and Szepesv \'a ri, C. (2020). Bandit algorithms https://tor-lattimore.com/downloads/book/book.pdf . Cambridge University Press
2020
-
[90]
Le Merrer, E., Pons, R., and Tr \'e dan, G. (2023). Algorithmic audits of algorithms, and the law. AI and Ethics , pages 1--11
2023
-
[91]
Liu, J., Nogueira, M., Fernandes, J., and Kantarci, B. (2021). Adversarial machine learning: A multilayer review of the state-of-the-art and challenges for wireless and mobile systems. IEEE Communications Surveys & Tutorials , 24(1):123--159
2021
-
[92]
T., Simchowitz, M., and Hardt, M
Liu, L. T., Simchowitz, M., and Hardt, M. (2019). The implicit fairness criterion of unconstrained learning. In International Conference on Machine Learning , pages 4051--4060. PMLR
2019
-
[93]
Liu, Q., Szepesv \'a ri, C., and Jin, C. (2022). Sample-efficient reinforcement learning of partially observable markov games. Advances in Neural Information Processing Systems , 35:18296--18308
2022
-
[94]
Liu, Y., Radanovic, G., Dimitrakakis, C., Mandal, D., and Parkes, D. C. (2017). Calibrated fairness in bandits
2017
-
[95]
K., Ramamurthy, K
Lohia, P. K., Ramamurthy, K. N., Bhide, M., Saha, D., Varshney, K. R., and Puri, R. (2019). Bias mitigation post-processing for individual and group fairness. In Icassp 2019-2019 ieee international conference on acoustics, speech and signal processing (icassp) , pages 2847--2851. IEEE
2019
-
[96]
T., Ruggieri, S., and Turini, F
Luong, B. T., Ruggieri, S., and Turini, F. (2011). k-nn as an implementation of situation testing for discrimination discovery and prevention. In Proceedings of the 17th ACM SIGKDD international conference on Knowledge discovery and data mining , pages 502--510
2011
-
[97]
Madiega, T. (2021). Artificial intelligence act. European Parliament: European Parliamentary Research Service
2021
-
[98]
Maneriker, P., Burley, C., and Parthasarathy, S. (2023). Online fairness auditing through iterative refinement. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , KDD '23, page 1665–1676, New York, NY, USA. Association for Computing Machinery
2023
-
[99]
Mangold, P., Perrot, M., Bellet, A., and Tommasi, M. (2023). Differential privacy has bounded impact on fairness in classification. In International Conference on Machine Learning , pages 23681--23705. PMLR
2023
-
[100]
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM computing surveys (CSUR) , 54(6):1--35
2021
-
[101]
Mohri, M. (2018). Foundations of machine learning
2018
-
[102]
and Shafer, J
Mutreja, S. and Shafer, J. (2023). Pac verification of statistical algorithms. In The Thirty Sixth Annual Conference on Learning Theory , pages 5021--5043. PMLR
2023
-
[103]
Novelli, C., Casolari, F., Rotolo, A., Taddeo, M., and Floridi, L. (2023). Taking ai risks seriously: a new assessment model for the ai act. AI & SOCIETY , pages 1--5
2023
-
[104]
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). Training language models to follow instructions with human feedback. Advances in neural information processing systems , 35:27730--27744
2022
-
[105]
Paltrinieri, N., Comfort, L., and Reniers, G. (2019). Learning about risk: Machine learning for risk assessment. Safety science , 118:475--486
2019
-
[106]
Pardau, S. L. (2018). The california consumer privacy act: Towards a european-style privacy regime in the united states. J. Tech. L. & Pol'y , 23:68
2018
-
[107]
Pentyala, S., Neophytou, N., Nascimento, A., De Cock, M., and Farnadi, G. (2022). Privfairfl: Privacy-preserving group fairness in federated learning. arXiv preprint arXiv:2205.11584
2022 arXiv
-
[108]
and Shmueli, E
Pessach, D. and Shmueli, E. (2022). A review on fairness in machine learning. ACM Computing Surveys (CSUR) , 55(3):1--44
2022
-
[109]
Pleiss, G., Raghavan, M., Wu, F., Kleinberg, J., and Weinberger, K. Q. (2017). On fairness and calibration. Advances in neural information processing systems , 30
2017
-
[110]
Salimi, B., Howe, B., and Suciu, D. (2019a). Data management for causal algorithmic fairness. arXiv preprint arXiv:1908.07924
1908 arXiv
-
[111]
Salimi, B., Rodriguez, L., Howe, B., and Suciu, D. (2019b). Capuchin: Causal database repair for algorithmic fairness. arXiv preprint arXiv:1902.08283
1902 arXiv
-
[112]
Salimi, B., Rodriguez, L., Howe, B., and Suciu, D. (2019c). Interventional fairness: Causal database repair for algorithmic fairness. In Proceedings of the 2019 International Conference on Management of Data , pages 793--810
2019
-
[113]
Shankar, S., Zamfirescu-Pereira, J., Hartmann, B., Parameswaran, A., and Arawjo, I. (2024). Who validates the validators? aligning llm-assisted evaluation of llm outputs with human preferences. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Tech...
2024
-
[114]
Shapley, L. S. (1953). Stochastic games. Proceedings of the national academy of sciences , 39(10):1095--1100
1953
-
[115]
and Basu, D
Shukla, A. and Basu, D. (2024). Preference-based Pure Exploration . In Advances in Neural Information Processing Systems (NeurIPS) , Vancouver (CA), Canada
2024
-
[116]
J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016). Mastering the game of go with deep neural networks and tree search. nature , 529(7587):484--489
2016
-
[117]
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2017). Mastering chess and shogi by self-play with a general reinforcement learning algorithm. arXiv preprint arXiv:1712.01815
2017 arXiv
-
[118]
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H. (2024). Preference ranking optimization for human alignment. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 38, pages 18990--18998
2024
-
[119]
Sorin, S. (1986). Asymptotic properties of a non-zero sum stochastic game. International Journal of Game Theory , 15:101--107
1986
-
[120]
Sun, B., Sun, J., Dai, T., and Zhang, L. (2021). Probabilistic verification of neural networks against group fairness
2021
-
[121]
Sutton, R. S. and Barto, A. G. (2018). Reinforcement learning: An introduction https://www.andrew.cmu.edu/user/rmorina/papers/SuttonBook.pdf . MIT press
2018
-
[122]
S., and Kizilcec, R
Tao, Y., Viberg, O., Baker, R. S., and Kizilcec, R. F. (2024). Cultural bias and cultural alignment of large language models . PNAS Nexus , 3(9):pgae346
2024
-
[123]
Tavara, S., Schliep, A., and Basu, D. (2021). https://drive.google.com/file/d/1ejIJsPboED_jDSI59RITSwXW78ZmzyX1/view Federated learning of oligonucleotide drug molecule thermodynamics with differentially private ADMM -based SVM . In Machine Learning and Principles and Practice...
2021
-
[124]
Vapnik, V. (1991). Principles of risk minimization for learning theory. Advances in neural information processing systems , 4
1991
-
[125]
and Zuiderveen Borgesius, F
Veale, M. and Zuiderveen Borgesius, F. (2021). Demystifying the draft eu artificial intelligence act—analysing the good, the bad, and the unclear elements of the proposed approach. Computer Law Review International , 22(4):97--112
2021
-
[126]
and Von dem Bussche, A
Voigt, P. and Von dem Bussche, A. (2017). The eu general data protection regulation (gdpr). A Practical Guide, 1st Ed., Cham: Springer International Publishing , 10(3152676):10--5555
2017
-
[127]
Wang, B., Gu, Q., Boedihardjo, M., Wang, L., Barekat, F., and Osher, S. J. (2020). Dp-lssgd: A stochastic optimization method to lift the utility in privacy-preserving erm. In Mathematical and Scientific Machine Learning , pages 328--351. PMLR
2020
-
[128]
Wang, Y., Zhong, W., Li, L., Mi, F., Zeng, X., Huang, W., Shang, L., Jiang, X., and Liu, Q. (2023). Aligning large language models with human: A survey. arXiv preprint arXiv:2307.12966
2023 arXiv
-
[129]
M., Vecchione, B., Qu, T., Cai, P., Smith, A., Investigators, C
Wright, L., Muenster, R. M., Vecchione, B., Qu, T., Cai, P., Smith, A., Investigators, C. . S., Metcalf, J., Matias, J. N., et al. (2024). Null compliance: Nyc local law 144 and the challenges of algorithm accountability. In The 2024 ACM Conference on Fairness, Accountability,...
2024
-
[130]
Xiao, J., Li, Z., Xie, X., Getzen, E., Fang, C., Long, Q., and Su, W. J. (2024). On the algorithmic bias of aligning large language models with rlhf: Preference collapse and matching regularization. arXiv preprint arXiv:2405.16455
2024 arXiv
-
[131]
and Zhang, C
Yan, T. and Zhang, C. (2022). Active fairness auditing. In International Conference on Machine Learning , pages 24929--24962. PMLR
2022
-
[132]
Yang, Y., Juntao, L., and Lingling, P. (2020). Multi-robot path planning based on a deep reinforcement learning dqn algorithm. CAAI Transactions on Intelligence Technology , 5(3):177--183
2020
-
[133]
Yu, T., Yao, Y., Zhang, H., He, T., Han, Y., Cui, G., Hu, J., Liu, Z., Zheng, H.-T., Sun, M., et al. (2024). Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. In Proceedings of the IEEE/CVF Conference on Computer Vision and...
2024
-
[134]
Zemel, R., Wu, Y., Swersky, K., Pitassi, T., and Dwork, C. (2013). Learning fair representations. In International conference on machine learning , pages 325--333. PMLR
2013
-
[135]
Zhang, K., Kakade, S., Basar, T., and Yang, L. (2020). Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity. Advances in Neural Information Processing Systems , 33:1166--1178
2020
-
[136]
Zhang, T., Zeng, Z., Xiao, Y., Zhuang, H., Chen, C., Foulds, J., and Pan, S. (2024). Genderalign: An alignment dataset for mitigating gender bias in large language models. arXiv preprint arXiv:2406.13925
2024 arXiv
-
[137]
M., Stiennon, N., Wu, J., Brown, T
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019). Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593
2019 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.