Pith. sign in

REVIEW 4 major objections 2 minor 137 references

The Fair Game: Auditing & Debiasing AI Algorithms Over Time

T0 review · 4 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper claims that fairness in machine learning can be kept adaptive by an RL loop where an auditor's changing bias definition automatically retunes a debiasing algorithm, without model redesign.

desk verdict The abstract proposes a genuinely interesting adjustable-auditor idea for fair ML, but the full text is an unrelated superconductivity paper, so there is nothing to review. read the letter →

arxiv 2508.06443 v1 pith:VTFX5AKO submitted 2025-08-08 cs.AI cs.CYcs.ETcs.GT

classification cs.AIcs.CYcs.ETcs.GT
keywords fairmachinelearningalgorithmicdebiasingreinforcementadaptivefairnessauditingevolvinggoalsbiasmetricsML
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes 'Fair Game', a framework that couples an auditor measuring bias in a machine-learning model's predictions with a reinforcement-learning debiasing algorithm, so the model's outputs are continuously adjusted to stay fair. Its central claim is that the fairness goal can be changed over time simply by changing the auditor's definition of bias, without retraining or redesigning the ML model. This matters because fairness standards evolve in society, while existing observational bias metrics conflict and often require ground truth or hindsight. The framework treats fairness as a moving target in an RL loop, simulating how ethical and legal frameworks evolve.

What carries the argument

The Auditor–Debiaser loop under reinforcement learning: the Auditor's bias measure serves as the reward signal for the RL debiaser, which alters the ML model's predictions. Because the reward is redefined by the Auditor, the fairness objective is non-stationary and can be swapped without touching the underlying model.

What would settle it

Simulate the loop with an auditor that alternates between two conflicting definitions, such as demographic parity and equalized odds, and measure the model's error and fairness against ground-truth labels; if the model tracks the auditor but ground-truth fairness degrades or the loop oscillates, the central claim of reliable adaptation fails.

Watch

Extended reading notes

Core claim

Fair Game is a two-component RL loop wrapped around any ML algorithm. An Auditor quantifies bias using a chosen observational definition and emits feedback; a Debiasing algorithm, driven by reinforcement learning, adjusts the ML predictions to minimize that defined unfairness. The author's key claim is that to change the fairness criterion, one only replaces or modifies the Auditor—the feedback it sends changes, and the RL debiaser adapts accordingly. Thus the same loop works pre- and post-deployment, and can track changing societal or legal fairness norms over time.

Load-bearing premise

The loop can only be fair if the auditor's observational bias measure is a reliable and sufficient signal for steering the model toward genuine fairness, yet the abstract itself notes such measures need ground truth or hindsight; the loop's adaptivity does not resolve that grounding gap.

Editorial extensions

If this is right

  • Deployed ML models can be kept fair as legal or ethical fairness criteria change, without model redesign.
  • Conflicting bias definitions can be handled by switching the auditor's metric rather than retraining the model.
  • Fairness is maintained both before and after deployment, addressing the retrospective limitation of observational measures.
  • The RL debiaser can track a non-stationary reward target, enabling simulation of evolving social norms.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The meaningfulness of the loop depends on the auditor's metric being a valid proxy for the fairness one actually cares about; an arbitrary auditor would define fairness circularly.
  • If the auditor's metric changes abruptly, the RL loop may become unstable; the paper does not address convergence under time-varying rewards.
  • The supplied full text is a physics manuscript on nickelate superconductors, not the described Fair Game framework, so the abstract's claims are not supported by derivations in the provided document.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 2 minor

Summary. The submitted manuscript, arXiv:2508.06443, titled 'The Fair Game: Auditing & Debiasing AI Algorithms Over Time,' proposes a framework in which an Auditor and a Debiasing algorithm are placed in a reinforcement learning loop around an ML algorithm. The abstract claims that this loop can 'assure fairness' and adapt predictions over time by modifying only the auditor, thereby simulating evolving ethical and legal frameworks. However, the provided full text is a condensed-matter paper on the multiorbital character of density waves in trilayer nickelate superconductors. It contains no mention of Fair Game, Auditor, Debiasing, reinforcement learning, fairness, or any component of the claimed framework. Thus the central claims are unsupported by any technical specification, derivation, or experimental validation.

Significance. If the proposed framework were operational, dynamic adaptation of fairness criteria via an auditor-debiaser RL loop would be a potentially significant step in fair machine learning. The abstract identifies a real limitation of observational bias definitions and proposes a plausible direction. However, the submitted manuscript provides no equations, no algorithm definition, no stability or convergence analysis, and no experiments. Because the full text is an unrelated superconductivity paper, the claimed contribution is entirely unverifiable. The paper therefore cannot be assessed as a scientific contribution in its current form.

major comments (4)
  1. [Full text] The full text of the manuscript is an unrelated condensed-matter paper on trilayer nickelate superconductors. It contains no discussion of Fair Game, the Auditor, the Debiasing algorithm, reinforcement learning, fairness metrics, or any of the components named in the abstract. This is a load-bearing defect: the central claim of an adaptive fairness loop has no supporting specification whatsoever. Neither a reader nor a referee can verify the proposed mechanism, and there is no basis for judging correctness, novelty, or reproducibility.
  2. [Abstract (observational bias limitations)] The abstract states that observational bias definitions 'can only be deployed if either the ground truth is known or only in retrospect,' yet the proposed loop feeds exactly such observational measures into the debiasing agent. The abstract does not explain how the auditor obtains ground truth, how it avoids the retrospective gap, or why optimizing against the auditor's own criteria constitutes genuine fairness rather than circular self-confirmation. This is not merely a presentation issue; it undermines the central claim that the framework can 'assure fairness' rather than internally defined compliance.
  3. [No algorithm specification or analysis] No equations, pseudocode, or formal description of the Fair Game loop are given. There is no specification of the state space, action space, reward signal, policy optimization method, convergence conditions, or stability under a changing reward target. Without these, the core claim that RL can adapt fairness goals over time 'by only modifying the auditor' cannot be evaluated, and the non-stationary reward problem introduced by a time-varying auditor is unaddressed.
  4. [No experiments or comparisons] The manuscript contains no empirical evaluation, no case studies, and no comparison to existing fair-ML or RL-for-fairness methods. The abstract's claims about pre- and post-deployment adaptation are therefore unsupported. This would be a major shortcoming even if the full text were on-topic, and it is decisive given the absence of any technical content.
minor comments (2)
  1. [Abstract] The abstract contains awkward repetitions, e.g., the name 'Fair Game' appearing twice in overlapping sentences, and undefined terms such as 'Auditor' and 'Debiasing algorithm' that are not described anywhere in the supplied text. These presentation issues are secondary to the substantive defects.
  2. [General] If the full-text mismatch is due to an administrative error, the authors should resubmit the correct manuscript with complete technical appendices. As it stands, the document is internally inconsistent: title and abstract do not correspond to the body.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation present: the full text is an unrelated superconductivity paper, and the abstract's Fair Game framework is asserted rather than derived.

full rationale

The provided full text is a condensed-matter paper on trilayer nickelate superconductors; it does not contain the Auditor, Debiasing algorithm, reinforcement-learning loop, or any equations or results belonging to the 'Fair Game' framework. There is therefore no derivation chain in which a claimed result can be shown to reduce to its inputs. The abstract's concession that observational bias definitions require ground truth or retrospective knowledge identifies a real limitation, and the proposed loop does not resolve that limitation, but that is a support/correctness gap rather than a circularity: the paper does not derive 'fairness' from the auditor's measure by equation or by construction. The statement that fairness goals can be adapted 'by only modifying the auditor' describes a feedback architecture, not a logical reduction of a conclusion to its premise. Without any actual derivation, no specific circular step can be exhibited, so the honest finding is no significant circularity.

Assumptions & free parameters 1 free parameters · 3 assumptions · 2 invented entities

For an abstract-only review, the ledger is inferred from the proposal. The framework's uncharged inputs are: freely chosen auditor criteria, stability of RL against a moving fairness target, sufficiency of observational bias measures (which the abstract itself identifies as problematic), and the equation of auditor updates with evolving ethical and legal frameworks. No datasets, external benchmarks, or fitted constants are mentioned, and the mismatched full text cannot be audited.

free parameters (1)
  • Auditor's bias criteria and their thresholds
    The abstract allows the auditor's fairness criteria to be modified over time but never specifies how they are set or updated; these choices determine what the loop optimizes and are freely chosen modeling inputs.
assumptions (3)
  • domain assumption The RL debiasing loop learns a stable policy while the auditor's reward-relevant criteria change over time
    The abstract proposes adapting fairness goals 'by only modifying the auditor' with no convergence or stability analysis for this moving-target reward setting.
  • domain assumption Observational bias measures quantified by the auditor are sufficient to drive genuine fairness improvement
    The abstract itself states observational definitions 'can only be deployed if either the ground truth is known or only in retrospect', yet the proposed loop feeds exactly such measures to the debiaser; their sufficiency is assumed, not shown.
  • domain assumption Modifying the auditor over time simulates the evolution of ethical and legal frameworks
    The abstract asserts this simulation goal with no model connecting auditor updates to actual social or legal change.
invented entities (2)
  • The Auditor
    purpose: Quantifies the biases of current concern and sends feedback to the debiasing agent
    The Auditor is the component that defines the fairness objective; no falsifiable handle outside the framework is described, so its outputs cannot be checked against an independent fairness benchmark.
  • The Debiasing algorithm (RL agent)
    purpose: Adjusts the ML algorithm's predictions using the auditor's feedback as a reward signal
    Second loop component; its behavior is entirely defined by the auditor's feedback, and no independent evaluation is described in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Fair Game: Auditing & Debiasing AI Algorithms Over Time." pith.science (2026). https://pith.science/paper/VTFX5AKO

@misc{pith2026250806443,
  author       = {Pith},
  title        = {Pith review of: The Fair Game: Auditing & Debiasing AI Algorithms Over Time},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VTFX5AKO}},
  note         = {Machine review of arXiv:2508.06443}
}
read the original abstract

An emerging field of AI, namely Fair Machine Learning (ML), aims to quantify different types of bias (also known as unfairness) exhibited in the predictions of ML algorithms, and to design new algorithms to mitigate them. Often, the definitions of bias used in the literature are observational, i.e. they use the input and output of a pre-trained algorithm to quantify a bias under concern. In reality,these definitions are often conflicting in nature and can only be deployed if either the ground truth is known or only in retrospect after deploying the algorithm. Thus,there is a gap between what we want Fair ML to achieve and what it does in a dynamic social environment. Hence, we propose an alternative dynamic mechanism,"Fair Game",to assure fairness in the predictions of an ML algorithm and to adapt its predictions as the society interacts with the algorithm over time. "Fair Game" puts together an Auditor and a Debiasing algorithm in a loop around an ML algorithm. The "Fair Game" puts these two components in a loop by leveraging Reinforcement Learning (RL). RL algorithms interact with an environment to take decisions, which yields new observations (also known as data/feedback) from the environment and in turn, adapts future decisions. RL is already used in algorithms with pre-fixed long-term fairness goals. "Fair Game" provides a unique framework where the fairness goals can be adapted over time by only modifying the auditor and the different biases it quantifies. Thus,"Fair Game" aims to simulate the evolution of ethical and legal frameworks in the society by creating an auditor which sends feedback to a debiasing algorithm deployed around an ML system. This allows us to develop a flexible and adaptive-over-time framework to build Fair ML systems pre- and post-deployment.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

137 extracted references · 66 canonical work pages

  1. [1]

    Agarwal, A., Beygelzimer, A., Dud \' k, M., Langford, J., and Wallach, H. (2018). A reductions approach to fair classification. In International conference on machine learning , pages 60--69. PMLR

  2. [2]

    Ajarra, A., Ghosh, B., and Basu, D. (2024). Active fourier auditor for estimating distributional properties of ml models. arXiv preprint arXiv:2410.08111

  3. [3]

    Albarghouthi, A., D'Antoni, L., Drews, S., and Nori, A. V. (2017). Fairsquare: probabilistic verification of program fairness. Proc. ACM Program. Lang. , 1(OOPSLA)

  4. [4]

    Altman, E. (1999). Constrained Markov decision processes https://hal.inria.fr/docs/00/07/41/09/PS/RR-2574.ps , volume 7. CRC Press

  5. [5]

    (23 May 2016)

    Angwin, J., Larson, J., Mattu, S., and Kirchner, L. (23 May 2016). Machine bias: There’s software used across the country to predict future criminals. And it’s biased against blacks, http://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing ProPublica

  6. [6]

    Annas, G. J. (2003). Hipaa regulations: a new era of medical-record privacy? New England Journal of Medicine , 348:1486

  7. [7]

    Arrow, K. (1971). The theory of discrimination

  8. [8]

    Assaf, D., Gutman, Y., Neuman, Y., Segal, G., Amit, S., Gefen-Halevi, S., Shilo, N., Epstein, A., Mor-Cohen, R., Biber, A., et al. (2020). Utilization of machine-learning models to accurately predict the risk for critical covid-19. Internal and emergency medicine , 15:1435--1443

Show all 137 references
  1. [9]

    and Basu, D

    Azize, A. and Basu, D. (2022). When Privacy Meets Partial Information: A Refined Analysis of Differentially Private Bandits . In Advances in Neural Information Processing Systems , New Orleans, United States

  2. [10]

    and Basu, D

    Azize, A. and Basu, D. (2024). Concentrated Differential Privacy for Bandits . In 2024 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages 78--109, Toronto, Canada. IEEE , IEEE

  3. [11]

    A., and Basu, D

    Azize, A., Jourdan, M., Marjani, A. A., and Basu, D. (2023). On the Complexity of Differentially Private Best-Arm Identification with Fixed Confidence . In NeurIPS 2023 -- Conference on Neural Information Processing Systems , volume 36, pages 71150--71194, New Orleans (US), Un...

  4. [12]

    Bagaric, M., Hunter, D., and Stobbs, N. (2019). Erasing the bias against using artificial intelligence to predict future criminality: algorithms are color blind and never tire. U. Cin. L. Rev. , 88:1037

  5. [13]

    Bai, Y., Jin, C., Mei, S., and Yu, T. (2022). Near-optimal learning of extensive-form games with imperfect information. In International Conference on Machine Learning , pages 1337--1382. PMLR

  6. [14]

    Barocas, S., Hardt, M., and Narayanan, A. (2023). Fairness and machine learning: Limitations and opportunities . MIT Press

  7. [15]

    Bastani, O., Zhang, X., and Solar-Lezama, A. (2019). Probabilistic verification of fairness properties via concentration. Proceedings of the ACM on Programming Languages , 3(OOPSLA):1--27

  8. [16]

    Basu, D., Dimitrakakis, C., and Tossou, A. (2020). Privacy in Multi-armed Bandits: Fundamental Definitions and Lower Bounds https://arxiv.org/pdf/1905.12298.pdf. NeurIPS Workshop on Privacy Preserving Machine Learning

  9. [17]

    Basu, D., Maillard, O.-A., and Mathieu, T. (2022). Bandits corrupted by Nature: Lower bounds on regret and robust optimistic algorithm https://arxiv.org/pdf/2203.03186.pdf. arXiv preprint arXiv:2203.03186 . (Submitted to COLT'22)

  10. [18]

    K., Dey, K., Hind, M., Hoffman, S

    Bellamy, R. K., Dey, K., Hind, M., Hoffman, S. C., Houde, S., Kannan, K., Lohia, P., Martino, J., Mehta, S., Mojsilovi \'c , A., et al. (2019). Ai fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias. IBM Journal of Research and Development , 63(4/5):4--1

  11. [19]

    Bellamy, R. K. E., Dey, K., Hind, M., Hoffman, S. C., Houde, S., Kannan, K., Lohia, P., Martino, J., Mehta, S., Mojsilovic, A., Nagar, S., Ramamurthy, K. N., Richards, J., Saha, D., Sattigeri, P., Singh, M., Varshney, K. R., and Zhang, Y. (2018). Ai fairness 360: An extensible...

  12. [20]

    Blumrosen, A. W. (1967). The duty of fair recruitment under the civil rights act of 1964. Rutgers L. Rev. , 22:465

  13. [21]

    Borca-Tasciuc, G., Guo, X., Bak, S., and Skiena, S. (2022). Provable fairness for neural network models using formal verification

  14. [22]

    Brown, G. W. (1951). Iterative solution of games by fictitious play. Act. Anal. Prod Allocation , 13(1):374

  15. [23]

    K., Segal, M., Basu, D., George, A.-M., and Dimitrakakis, C

    Buening, T. K., Segal, M., Basu, D., George, A.-M., and Dimitrakakis, C. (2022). On Meritocracy in Optimal Set Selection . In EAAMO 2022- Equity and Access in Algorithms, Mechanisms, and Optimization , Arlington, United States. ACM

  16. [24]

    Bénesse, C., Gamboa, F., Loubes, J.-M., and Boissin, T. (2021). Fairness seen as global sensitivity analysis

  17. [25]

    and Z liobait \.e , I

    Calders, T. and Z liobait \.e , I. (2013). Why unbiased computational processes can lead to discriminative decision procedures. In Discrimination and Privacy in the Information Society: Data mining and profiling in large databases , pages 43--57. Springer

  18. [26]

    D., and Dubhashi, D

    Carlsson, E., Basu, D., Johansson, F. D., and Dubhashi, D. (2024). Pure Exploration in Bandits with Linear Constraints . In International Conference on Artificial Intelligence and Statistics , volume 238 of Proceedings of Machine Learning Research (PMLR) , pages 334--342, Vale...

  19. [27]

    cas, S., Hardt, M., and Narayanan, A. (2023). Fairness and Machine Learning: Limitations and Opportunities . MIT Press

  20. [28]

    and Haas, C

    Caton, S. and Haas, C. (2024). Fairness in machine learning: A survey. ACM Computing Surveys , 56(7):1--38

  21. [29]

    E., Huang, L., Keswani, V., and Vishnoi, N

    Celis, L. E., Huang, L., Keswani, V., and Vishnoi, N. K. (2019). Classification with fairness constraints: A meta-algorithm with provable guarantees. In Proceedings of the conference on fairness, accountability, and transparency , pages 319--328

  22. [30]

    Chaudhuri, K., Monteleoni, C., and Sarwate, A. D. (2011). Differentially private empirical risk minimization. Journal of Machine Learning Research , 12(3)

  23. [31]

    R., and Liu, H

    Cheng, L., Varshney, K. R., and Liu, H. (2021). Socially responsible AI algorithms: Issues, purposes, and challenges. Journal of Artificial Intelligence Research , 71:1137--1181

  24. [32]

    Cherian, J. J. and Cand \`e s, E. J. (2024). Statistical inference for fairness auditing. Journal of Machine Learning Research , 25(149):1--49

  25. [33]

    Chierichetti, F., Kumar, R., Lattanzi, S., and Vassilvtiskii, S. (2019). Matroids, matchings, and fairness. In The 22nd international conference on artificial intelligence and statistics , pages 2212--2220. PMLR

  26. [34]

    Chouldechova, A. (2017). Fair prediction with disparate impact: A study of bias in recidivism prediction instruments. Big data , 5(2):153--163

  27. [35]

    and Roth, A

    Chouldechova, A. and Roth, A. (2018). The frontiers of fairness in machine learning. arXiv preprint arXiv:1810.08810

  28. [36]

    and Roth, A

    Chouldechova, A. and Roth, A. (2020). A snapshot of the frontiers of fairness in machine learning https://dl.acm.org/doi/pdf/10.1145/3376898. Communications of the ACM , 63(5):82--89

  29. [37]

    F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

    Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in neural information processing systems , 30

  30. [38]

    Chzhen, E., Denis, C., Hebiri, M., Oneto, L., and Pontil, M. (2020). Fair regression with Wasserstein barycenters https://arxiv.org/pdf/2006.07286.pdf. Advances in Neural Information Processing Systems , 33:7321--7331

  31. [39]

    H., Jacobs, B

    Conitzer, V., Freedman, R., Heitzig, J., Holliday, W. H., Jacobs, B. M., Lambert, N., Moss \'e , M., Pacuit, E., Russell, S., Schoelkopf, H., et al. (2024). Position: Social choice should guide ai alignment in dealing with diverse human feedback. In Forty-first International C...

  32. [40]

    Corbett-Davies, S., Pierson, E., Feller, A., Goel, S., and Huq, A. (2017). Algorithmic decision making and the cost of fairness. In Proceedings of the 23rd acm sigkdd international conference on knowledge discovery and data mining , pages 797--806

  33. [41]

    Cotter, A., Jiang, H., Gupta, M., Wang, S., Narayan, T., You, S., and Sridharan, K. (2019). Optimization with non-differentiable constraints with applications to fairness, recall, churn, and other goals. Journal of Machine Learning Research , 20(172):1--59

  34. [42]

    D a browski, . D. and Suska, M. (2022). The European Union Digital Single Market: Europe's Digital Transformation . Routledge

  35. [43]

    Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y. (2023). Safe rlhf: Safe reinforcement learning from human feedback. arXiv preprint arXiv:2310.12773

  36. [44]

    Dandekar, A., Basu, D., and Bressan, S. (2018). Differential privacy for regularised linear regression. In International Conference on Database and Expert Systems Applications , pages 483--491. Springer

  37. [45]

    and Basu, D

    Das, U. and Basu, D. (2024). Learning to explore with lagrangians for bandits under unknown constraints. In Seventeenth European Workshop on Reinforcement Learning

  38. [46]

    Daskalakis, C., Golowich, N., and Zhang, K. (2023). The complexity of markov equilibrium in stochastic games. In The Thirty Sixth Annual Conference on Learning Theory , pages 4180--4234. PMLR

  39. [47]

    Devroye, L., Gy \"o rfi, L., and Lugosi, G. (2013). A probabilistic theory of pattern recognition , volume 31. Springer Science & Business Media

  40. [48]

    and Roth, A

    Dwork, C. and Roth, A. (2014). The algorithmic foundations of differential privacy https://www.tau.ac.il/ saharon/BigData2018/privacybook.pdf. Found. Trends Theor. Comput. Sci. , 9(3-4):211--407

  41. [49]

    Elie, R., Perolat, J., Lauri \`e re, M., Geist, M., and Pietquin, O. (2020). On the convergence of model free learning in mean field games. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 34, pages 7143--7150

  42. [50]

    Eriksson, H., Basu, D., Alibeigi, M., and Dimitrakakis, C. (2022a). Risk-Sensitive Bayesian Games for Multi-Agent Reinforcement Learning under Policy Uncertainty . In OptLearnMAS@AAMAS , Workshop on Optimization and Learning in Multiagent Systems at International Conference on...

  43. [51]

    Eriksson, H., Basu, D., Alibeigi, M., and Dimitrakakis, C. (2022b). SENTINEL: Taming Uncertainty with Ensemble-based Distributional Reinforcement Learning . In UAI 2022- Proceedings of the Thirty-Eighth Conference on Uncertainty in Artificial Intelligence , volume 180 of Proce...

  44. [52]

    A., Moeller, J., Scheidegger, C., and Venkatasubramanian, S

    Feldman, M., Friedler, S. A., Moeller, J., Scheidegger, C., and Venkatasubramanian, S. (2015). Certifying and removing disparate impact. In proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining , pages 259--268

  45. [53]

    Feldman, V., Guruswami, V., Raghavendra, P., and Wu, Y. (2012). Agnostic learning of monomials by halfspaces is hard. SIAM Journal on Computing , 41(6):1558--1590

  46. [54]

    Fiegel, C., M \'e nard, P., Kozuno, T., Munos, R., Perchet, V., and Valko, M. (2023). Adapting to game trees in zero-sum imperfect information games. In International Conference on Machine Learning , pages 10093--10135. PMLR

  47. [55]

    Fiss, O. M. (1970). A theory of fair employment laws. U. Chi. L. Rev. , 38:235

  48. [56]

    and Basu, D

    Flet-Berliac, Y. and Basu, D. (2022). SAAC: Safe Reinforcement Learning as an Adversarial Game of Actor-Critics . In RLDM 2022 - The Multi-disciplinary Conference on Reinforcement Learning and Decision Making , Providence, United States. Accepted at the 5th Multi-disciplinary ...

  49. [57]

    Gajane, P., Saxena, A., Tavakol, M., Fletcher, G., and Pechenizkiy, M. (2022). Survey on fair reinforcement learning: Theory and practice. arXiv preprint arXiv:2205.10032

  50. [58]

    Galhotra, S., Brun, Y., and Meliou, A. (2017). Fairness testing: testing software for discrimination. In Proceedings of the 2017 11th Joint meeting on foundations of software engineering , pages 498--510

  51. [59]

    and Tamayo, P

    Galindo, J. and Tamayo, P. (2000). Credit risk assessment using statistical and machine learning: basic methodology and risk modeling applications. Computational economics , 15:107--143

  52. [60]

    Ghosh, B., Basu, D., and Meel, K. (2023a). ''How Biased are Your Features?'': Computing Fairness Influence Functions with Global Sensitivity Analysis . In FAccT '23: the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages 138--148, Chicago IL, United States. ACM

  53. [61]

    Ghosh, B., Basu, D., and Meel, K. S. (2021). Justicia: A stochastic sat approach to formally verify fairness. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 35, pages 7554--7563

  54. [62]

    Ghosh, B., Basu, D., and Meel, K. S. (2022a). Algorithmic fairness verification with graphical models. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 36, pages 9539--9548

  55. [63]

    Ghosh, B., Basu, D., and Meel, K. S. (2022b). Algorithmic fairness verification with graphical models. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 36, pages 9539--9548

  56. [64]

    Ghosh, B., Basu, D., and Meel, K. S. (2022c). Algorithmic fairness verification with graphical models . In AAAI-2022 - 36th AAAI Conference on Artificial Intelligence , volume 2, Virtual, United States

  57. [65]

    Ghosh, B., Basu, D., and Meel, K. S. (2023b). ``how biased are your features?”: Computing fairness influence functions with global sensitivity analysis. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages 138--148

  58. [66]

    Giannou, A., Lotidis, K., Mertikopoulos, P., and Vlatakis-Gkaragkounis, E.-V. (2022). On the convergence of policy gradient methods to nash equilibria in general stochastic games. Advances in Neural Information Processing Systems , 35:7128--7141

  59. [67]

    Godinot, A., Le Merrer, E., Tr \'e dan, G., Penzo, C., and Ta \"i ani, F. (2023). Change-Relaxed Active Fairness Auditing . In RJCIA 2023 - 21e Rencontres des Jeunes Chercheurs en Intelligence Artificiel , CNIA, pages 91--96, Strasbourg, France. Association Fran c aise pour l'...

  60. [68]

    Godinot, A., Le Merrer, E., Tr \'e dan, G., Penzo, C., and Ta \" ani, F. (2024). Under manipulations, are some ai models harder to audit? In 2024 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages 644--664. IEEE

  61. [69]

    N., Shafer, J., and Yehudayoff, A

    Goldwasser, S., Rothblum, G. N., Shafer, J., and Yehudayoff, A. (2021). Interactive proofs for verifying machine learning. In 12th Innovations in Theoretical Computer Science Conference (ITCS 2021) . Schloss-Dagstuhl-Leibniz Zentrum f \"u r Informatik

  62. [70]

    Gordaliza, P., Del Barrio, E., Fabrice, G., and Loubes, J.-M. (2019). Obtaining fairness using optimal transport theory. In International conference on machine learning , pages 2357--2365. PMLR

  63. [71]

    Gorti, A., Gaur, M., and Chadha, A. (2024). Unboxing occupational bias: Grounded debiasing llms with us labor data. arXiv preprint arXiv:2408.11247

  64. [72]

    Gu, S., Holly, E., Lillicrap, T., and Levine, S. (2017). Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In 2017 IEEE international conference on robotics and automation (ICRA) , pages 3389--3396. IEEE

  65. [73]

    Gy \"o rfi, L., Kohler, M., Krzyzak, A., and Walk, H. (2006). A distribution-free theory of nonparametric regression . Springer Science & Business Media

  66. [74]

    and Domingo-Ferrer, J

    Hajian, S. and Domingo-Ferrer, J. (2012). A methodology for direct and indirect discrimination prevention in data mining. IEEE transactions on knowledge and data engineering , 25(7):1445--1459

  67. [75]

    Hardt, M., Price, E., and Srebro, N. (2016). Equality of opportunity in supervised learning. Advances in neural information processing systems , 29

  68. [76]

    H \'e bert-Johnson, U., Kim, M., Reingold, O., and Rothblum, G. (2018). Multicalibration: Calibration for the (computationally-identifiable) masses. In International Conference on Machine Learning , pages 1939--1948. PMLR

  69. [77]

    and Krause, A

    Heidari, H. and Krause, A. (2018). Preventing disparate treatment in sequential decision making. In IJCAI , pages 2248--2254

  70. [78]

    M., Harman, M., and Sarro, F

    Hort, M., Chen, Z., Zhang, J. M., Harman, M., and Sarro, F. (2024). Bias mitigation for machine learning classifiers: A comprehensive survey. ACM Journal on Responsible Computing , 1(2):1--52

  71. [79]

    Huber, P. J. (1981). Robust statistics. Wiley Series in Probability and Mathematical Statistics

  72. [80]

    Janaro, C. (2023). Nyc local law 144: A failed attempt at regulating ai in hiring

  73. [81]

    Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y. (2024). Beavertails: Towards improved safety alignment of llm via a human-preference dataset. Advances in Neural Information Processing Systems , 36

  74. [82]

    and Nachum, O

    Jiang, H. and Nachum, O. (2020). Identifying and correcting label bias in machine learning. In International conference on artificial intelligence and statistics , pages 702--712. PMLR

  75. [83]

    and Calders, T

    Kamiran, F. and Calders, T. (2012). Data preprocessing techniques for classification without discrimination. Knowledge and information systems , 33(1):1--33

  76. [84]

    Kamishima, T., Akaho, S., Asoh, H., and Sakuma, J. (2012). Fairness-aware classifier with prejudice remover regularizer. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2012, Bristol, UK, September 24-28, 2012. Proceedings, Part II 23 ,...

  77. [85]

    Kidder, W. C. (2001). Does the lsat mirror or magnify racial and ethnic differences in educational attainment: A study of equally achieving elite college students. Calif. L. Rev. , 89:1055

  78. [86]

    Kim, M., Reingold, O., and Rothblum, G. (2018). Fairness through computationally-bounded awareness. Advances in neural information processing systems , 31

  79. [87]

    Kleinberg, J., Mullainathan, S., and Raghavan, M. (2016). Inherent trade-offs in the fair determination of risk scores. arXiv preprint arXiv:1609.05807

  80. [88]

    Kleine Buening, T., Segal, M., Basu, D., George, A.-M., and Dimitrakakis, C. (2022). On meritocracy in optimal set selection. In Equity and Access in Algorithms, Mechanisms, and Optimization , pages 1--14. ACM

  81. [89]

    and Szepesv \'a ri, C

    Lattimore, T. and Szepesv \'a ri, C. (2020). Bandit algorithms https://tor-lattimore.com/downloads/book/book.pdf . Cambridge University Press

  82. [90]

    Le Merrer, E., Pons, R., and Tr \'e dan, G. (2023). Algorithmic audits of algorithms, and the law. AI and Ethics , pages 1--11

  83. [91]

    Liu, J., Nogueira, M., Fernandes, J., and Kantarci, B. (2021). Adversarial machine learning: A multilayer review of the state-of-the-art and challenges for wireless and mobile systems. IEEE Communications Surveys & Tutorials , 24(1):123--159

  84. [92]

    T., Simchowitz, M., and Hardt, M

    Liu, L. T., Simchowitz, M., and Hardt, M. (2019). The implicit fairness criterion of unconstrained learning. In International Conference on Machine Learning , pages 4051--4060. PMLR

  85. [93]

    Liu, Q., Szepesv \'a ri, C., and Jin, C. (2022). Sample-efficient reinforcement learning of partially observable markov games. Advances in Neural Information Processing Systems , 35:18296--18308

  86. [94]

    Liu, Y., Radanovic, G., Dimitrakakis, C., Mandal, D., and Parkes, D. C. (2017). Calibrated fairness in bandits

  87. [95]

    K., Ramamurthy, K

    Lohia, P. K., Ramamurthy, K. N., Bhide, M., Saha, D., Varshney, K. R., and Puri, R. (2019). Bias mitigation post-processing for individual and group fairness. In Icassp 2019-2019 ieee international conference on acoustics, speech and signal processing (icassp) , pages 2847--2851. IEEE

  88. [96]

    T., Ruggieri, S., and Turini, F

    Luong, B. T., Ruggieri, S., and Turini, F. (2011). k-nn as an implementation of situation testing for discrimination discovery and prevention. In Proceedings of the 17th ACM SIGKDD international conference on Knowledge discovery and data mining , pages 502--510

  89. [97]

    Madiega, T. (2021). Artificial intelligence act. European Parliament: European Parliamentary Research Service

  90. [98]

    Maneriker, P., Burley, C., and Parthasarathy, S. (2023). Online fairness auditing through iterative refinement. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , KDD '23, page 1665–1676, New York, NY, USA. Association for Computing Machinery

  91. [99]

    Mangold, P., Perrot, M., Bellet, A., and Tommasi, M. (2023). Differential privacy has bounded impact on fairness in classification. In International Conference on Machine Learning , pages 23681--23705. PMLR

  92. [100]

    Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM computing surveys (CSUR) , 54(6):1--35

  93. [101]

    Mohri, M. (2018). Foundations of machine learning

  94. [102]

    and Shafer, J

    Mutreja, S. and Shafer, J. (2023). Pac verification of statistical algorithms. In The Thirty Sixth Annual Conference on Learning Theory , pages 5021--5043. PMLR

  95. [103]

    Novelli, C., Casolari, F., Rotolo, A., Taddeo, M., and Floridi, L. (2023). Taking ai risks seriously: a new assessment model for the ai act. AI & SOCIETY , pages 1--5

  96. [104]

    Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). Training language models to follow instructions with human feedback. Advances in neural information processing systems , 35:27730--27744

  97. [105]

    Paltrinieri, N., Comfort, L., and Reniers, G. (2019). Learning about risk: Machine learning for risk assessment. Safety science , 118:475--486

  98. [106]

    Pardau, S. L. (2018). The california consumer privacy act: Towards a european-style privacy regime in the united states. J. Tech. L. & Pol'y , 23:68

  99. [107]

    Pentyala, S., Neophytou, N., Nascimento, A., De Cock, M., and Farnadi, G. (2022). Privfairfl: Privacy-preserving group fairness in federated learning. arXiv preprint arXiv:2205.11584

  100. [108]

    and Shmueli, E

    Pessach, D. and Shmueli, E. (2022). A review on fairness in machine learning. ACM Computing Surveys (CSUR) , 55(3):1--44

  101. [109]

    Pleiss, G., Raghavan, M., Wu, F., Kleinberg, J., and Weinberger, K. Q. (2017). On fairness and calibration. Advances in neural information processing systems , 30

  102. [110]

    Salimi, B., Howe, B., and Suciu, D. (2019a). Data management for causal algorithmic fairness. arXiv preprint arXiv:1908.07924

  103. [111]

    Salimi, B., Rodriguez, L., Howe, B., and Suciu, D. (2019b). Capuchin: Causal database repair for algorithmic fairness. arXiv preprint arXiv:1902.08283

  104. [112]

    Salimi, B., Rodriguez, L., Howe, B., and Suciu, D. (2019c). Interventional fairness: Causal database repair for algorithmic fairness. In Proceedings of the 2019 International Conference on Management of Data , pages 793--810

  105. [113]

    Shankar, S., Zamfirescu-Pereira, J., Hartmann, B., Parameswaran, A., and Arawjo, I. (2024). Who validates the validators? aligning llm-assisted evaluation of llm outputs with human preferences. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Tech...

  106. [114]

    Shapley, L. S. (1953). Stochastic games. Proceedings of the national academy of sciences , 39(10):1095--1100

  107. [115]

    and Basu, D

    Shukla, A. and Basu, D. (2024). Preference-based Pure Exploration . In Advances in Neural Information Processing Systems (NeurIPS) , Vancouver (CA), Canada

  108. [116]

    J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al

    Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016). Mastering the game of go with deep neural networks and tree search. nature , 529(7587):484--489

  109. [117]

    Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2017). Mastering chess and shogi by self-play with a general reinforcement learning algorithm. arXiv preprint arXiv:1712.01815

  110. [118]

    Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H. (2024). Preference ranking optimization for human alignment. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 38, pages 18990--18998

  111. [119]

    Sorin, S. (1986). Asymptotic properties of a non-zero sum stochastic game. International Journal of Game Theory , 15:101--107

  112. [120]

    Sun, B., Sun, J., Dai, T., and Zhang, L. (2021). Probabilistic verification of neural networks against group fairness

  113. [121]

    Sutton, R. S. and Barto, A. G. (2018). Reinforcement learning: An introduction https://www.andrew.cmu.edu/user/rmorina/papers/SuttonBook.pdf . MIT press

  114. [122]

    S., and Kizilcec, R

    Tao, Y., Viberg, O., Baker, R. S., and Kizilcec, R. F. (2024). Cultural bias and cultural alignment of large language models . PNAS Nexus , 3(9):pgae346

  115. [123]

    Tavara, S., Schliep, A., and Basu, D. (2021). https://drive.google.com/file/d/1ejIJsPboED_jDSI59RITSwXW78ZmzyX1/view Federated learning of oligonucleotide drug molecule thermodynamics with differentially private ADMM -based SVM . In Machine Learning and Principles and Practice...

  116. [124]

    Vapnik, V. (1991). Principles of risk minimization for learning theory. Advances in neural information processing systems , 4

  117. [125]

    and Zuiderveen Borgesius, F

    Veale, M. and Zuiderveen Borgesius, F. (2021). Demystifying the draft eu artificial intelligence act—analysing the good, the bad, and the unclear elements of the proposed approach. Computer Law Review International , 22(4):97--112

  118. [126]

    and Von dem Bussche, A

    Voigt, P. and Von dem Bussche, A. (2017). The eu general data protection regulation (gdpr). A Practical Guide, 1st Ed., Cham: Springer International Publishing , 10(3152676):10--5555

  119. [127]

    Wang, B., Gu, Q., Boedihardjo, M., Wang, L., Barekat, F., and Osher, S. J. (2020). Dp-lssgd: A stochastic optimization method to lift the utility in privacy-preserving erm. In Mathematical and Scientific Machine Learning , pages 328--351. PMLR

  120. [128]

    Wang, Y., Zhong, W., Li, L., Mi, F., Zeng, X., Huang, W., Shang, L., Jiang, X., and Liu, Q. (2023). Aligning large language models with human: A survey. arXiv preprint arXiv:2307.12966

  121. [129]

    M., Vecchione, B., Qu, T., Cai, P., Smith, A., Investigators, C

    Wright, L., Muenster, R. M., Vecchione, B., Qu, T., Cai, P., Smith, A., Investigators, C. . S., Metcalf, J., Matias, J. N., et al. (2024). Null compliance: Nyc local law 144 and the challenges of algorithm accountability. In The 2024 ACM Conference on Fairness, Accountability,...

  122. [130]

    Xiao, J., Li, Z., Xie, X., Getzen, E., Fang, C., Long, Q., and Su, W. J. (2024). On the algorithmic bias of aligning large language models with rlhf: Preference collapse and matching regularization. arXiv preprint arXiv:2405.16455

  123. [131]

    and Zhang, C

    Yan, T. and Zhang, C. (2022). Active fairness auditing. In International Conference on Machine Learning , pages 24929--24962. PMLR

  124. [132]

    Yang, Y., Juntao, L., and Lingling, P. (2020). Multi-robot path planning based on a deep reinforcement learning dqn algorithm. CAAI Transactions on Intelligence Technology , 5(3):177--183

  125. [133]

    Yu, T., Yao, Y., Zhang, H., He, T., Han, Y., Cui, G., Hu, J., Liu, Z., Zheng, H.-T., Sun, M., et al. (2024). Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. In Proceedings of the IEEE/CVF Conference on Computer Vision and...

  126. [134]

    Zemel, R., Wu, Y., Swersky, K., Pitassi, T., and Dwork, C. (2013). Learning fair representations. In International conference on machine learning , pages 325--333. PMLR

  127. [135]

    Zhang, K., Kakade, S., Basar, T., and Yang, L. (2020). Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity. Advances in Neural Information Processing Systems , 33:1166--1178

  128. [136]

    Zhang, T., Zeng, Z., Xiao, Y., Zhuang, H., Chen, C., Foulds, J., and Pan, S. (2024). Genderalign: An alignment dataset for mitigating gender bias in large language models. arXiv preprint arXiv:2406.13925

  129. [137]

    M., Stiennon, N., Wu, J., Brown, T

    Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019). Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.