REVIEW 2 major objections 1 minor 52 references
Positive and Negative Determinant Strategies in Repeated Games with Behavior-Value Inconsistency
T0 review · 2 major / 1 minor · reviewed 2026-07-02 · grok-4.3
Pith's one-line read Behavior-value inconsistency eliminates zero-determinant strategies but enables positive and negative determinant strategies for unilateral payoff control.
desk verdict The paper adds an internal cost for behavior-value inconsistency in repeated games, claims this eliminates ZD strategies, and introduces positive/negative determinant strategies that enforce affine payoff combinations instead. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Positive/negative determinant strategy, which forces an affine combination of the two players' average payoffs to be positive or negative.
What would settle it
A calculation or simulation showing that a zero-determinant strategy still succeeds in enforcing its payoff relation, or that a positive/negative determinant strategy fails to enforce the claimed affine combination, once the inconsistency cost is subtracted from the payoffs.
Extended reading notes
Core claim
We prove that ZD strategy does not exist if the cost via behavior-value inconsistency is present. Instead, we find a new class of repeated strategies that enforce a unilateral payoff control, which is termed as positive/negative determinant strategy. The found strategy allows an individual to enforce an affine combination of two individuals' average payoffs above/below zero. Consequently, a focal individual is able to unilaterally control the opponent's payoff below a given value via negative determinant strategy, and a focal individual is able to get more payoff than the opponent via positive determinant strategy. We also find that the control ability of positive/negative determinant strate
Load-bearing premise
An agent pays an internal cost exactly when its behavior differs from its internal thought or value, and this cost structure is enough to remove ZD strategies while permitting the new determinant strategies.
Editorial extensions
If this is right
- Zero-determinant strategies cannot be used once inconsistency costs are included in the payoff structure.
- A focal player using a negative determinant strategy can unilaterally force the opponent's average payoff below any chosen level.
- A focal player using a positive determinant strategy can unilaterally ensure its own average payoff exceeds the opponent's.
- The payoff control achieved by positive/negative determinant strategies is strictly stronger than the control achieved by zero-determinant strategies.
Reading between the lines
- Models of reciprocity that assume no internal costs may systematically overstate the reach of zero-determinant strategies.
- Artificial agents could adopt positive/negative determinant strategies to achieve similar unilateral controls while remaining internally consistent.
- Empirical tests could check whether human subjects in repeated games behave as if they incur costs for action-value mismatches.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a repeated-game framework incorporating an internal cost for behavior-value inconsistency. It claims to prove that zero-determinant (ZD) strategies cease to exist under this cost model. In their place, the authors define a new class of 'positive/negative determinant strategies' that enforce an affine combination of the two players' average payoffs to lie strictly above or below zero. These strategies are asserted to permit unilateral control of the opponent's payoff (negative determinant) or to guarantee a higher payoff than the opponent (positive determinant), with superior control ability relative to classic ZD strategies.
Significance. If the claimed non-existence result and the construction of the new determinant strategies hold, the work would usefully extend the theory of unilateral payoff control by incorporating an internal consistency cost that is relevant to models of intelligent agents. The paper would thereby highlight a limitation of standard ZD analyses that omit such costs. No machine-checked proofs, reproducible code, or parameter-free derivations are described.
major comments (2)
- The abstract and reader's summary assert proofs of ZD non-existence and of the control properties of positive/negative determinant strategies, yet no derivations, payoff matrices, recurrence relations, or explicit strategy definitions appear in the provided text. Without these, the central claims cannot be verified and the soundness assessment remains low.
- The framework defines the new strategies relative to a specific inconsistency-cost model. It is unclear whether the claimed affine-control property follows independently of the particular functional form chosen for the cost or whether it reduces to a fitted parameter (cf. the reader's note on free_parameters = ['inconsistency cost']).
minor comments (1)
- The abstract contains minor grammatical issues ('provided what agents do differ', 'termed as') that should be corrected for clarity.
Simulated Author's Rebuttal
We thank the referee for the detailed report and the opportunity to clarify our contributions. We address the two major comments point by point below. The full manuscript contains the requested derivations, but we will revise to make them more prominent and to address generality concerns.
read point-by-point responses
-
Referee: The abstract and reader's summary assert proofs of ZD non-existence and of the control properties of positive/negative determinant strategies, yet no derivations, payoff matrices, recurrence relations, or explicit strategy definitions appear in the provided text. Without these, the central claims cannot be verified and the soundness assessment remains low.
Authors: The full manuscript (Sections 3 and 4) contains the non-existence proof for ZD strategies (Theorem 1), explicit strategy definitions for positive/negative determinant strategies, payoff matrices for the iterated Prisoner's Dilemma, and the recurrence relations used to derive the affine control property. We will revise the submission to include a dedicated appendix with all derivations, matrices, and strategy update rules to facilitate verification. revision: yes
-
Referee: The framework defines the new strategies relative to a specific inconsistency-cost model. It is unclear whether the claimed affine-control property follows independently of the particular functional form chosen for the cost or whether it reduces to a fitted parameter (cf. the reader's note on free_parameters = ['inconsistency cost']).
Authors: The non-existence of ZD strategies holds for any strictly positive inconsistency cost (i.e., whenever behavior differs from internal value). The affine-control property of positive/negative determinant strategies is derived under this general condition and does not depend on a specific functional form; the linear cost used in the paper is for concrete illustration only. The inconsistency cost is an explicit model parameter, not a fitted one. We will add a new subsection clarifying the generality of the results and the role of the cost parameter. revision: partial
Circularity Check
No significant circularity detected
full rationale
The paper introduces an explicit new cost model for behavior-value inconsistency, proves non-existence of ZD strategies within that model via direct analysis of the modified payoff structure, and defines positive/negative determinant strategies as the resulting class that enforces the claimed affine combinations. No step reduces by construction to a fitted parameter renamed as prediction, a self-citation chain, or a self-definitional loop; the derivations rest on the introduced framework and standard repeated-game algebra rather than importing uniqueness or ansatzes from prior author work. The central claims remain independent of the inputs.
Assumptions & free parameters
free parameters (1)
- inconsistency cost
assumptions (1)
- domain assumption An individual pays the internal cost if the behavior is inconsistent with the internal thought.
invented entities (1)
-
positive/negative determinant strategy
Cite this review
Pith. "Pith review of Positive and Negative Determinant Strategies in Repeated Games with Behavior-Value Inconsistency." pith.science (2026). https://pith.science/paper/KFPTZO5A
@misc{pith2026260700625,
author = {Pith},
title = {Pith review of: Positive and Negative Determinant Strategies in Repeated Games with Behavior-Value Inconsistency},
year = {2026},
howpublished = {\url{https://pith.science/paper/KFPTZO5A}},
note = {Machine review of arXiv:2607.00625}
}
read the original abstract
Direct reciprocity, based on the repeated interactions, is a fundamental mechanism to promote cooperation. Zero-determinant (ZD) strategies have opened an avenue for unilateral payoff control. However, previous studies neglect internal costs provided what agents do differ from what agents think, which is crucial for decision making of intelligent agents. Motivated by this, we establish a game theoretical framework by assuming that an individual pays the internal cost if the behavior is inconsistent with the internal thought. We prove that ZD strategy does not exist if the cost via behavior-value inconsistency is present. Instead, we find a new class of repeated strategies that enforce a unilateral payoff control, which is termed as positive/negative determinant strategy. The found strategy allows an individual to enforce an affine combination of two individuals' average payoffs above/below zero. Consequently, a focal individual is able to unilaterally control the opponent's payoff below a given value via negative determinant strategy, and a focal individual is able to get more payoff than the opponent via positive determinant strategy. We also find that the control ability of positive/negative determinant strategies is better off than that of ZD strategies. Our work highlights the importance of inconsistency between the behavior and value on payoff control, which is typically absent in classic ZD strategies.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Christoph Adami and Arend Hintze. Evolutionary instability of zero-determinant strategies demonstrates that winning is not everything.Nature communications, 4(1):2193, 2013
work page 2013
-
[2]
The iterated prisoner’s dilemma: good strategies and their dynamics
Ethan Akin. The iterated prisoner’s dilemma: good strategies and their dynamics. Ergodic Theory, Advances in Dynamical Systems, pages 77–107, 2016
work page 2016
-
[3]
Marco Archetti and Istvan Scheuring. Game theory of public goods in one-shot social dilemmas without assortment.Journal of theoretical biology, 299:9–20, 2012
work page 2012
-
[4]
Studies of independence and conformity: I
Solomon E Asch. Studies of independence and conformity: I. a minority of one against a unanimous majority.Psychological monographs: General and applied, 70(9):1, 1956
work page 1956
-
[5]
Gustavo Assun¸ c˜ ao, Bruno Patr˜ ao, Miguel Castelo-Branco, and Paulo Menezes. An overview of emotion in artificial intelligence.IEEE Transactions on Artificial Intel- ligence, 3(6):867–886, 2022
work page 2022
-
[6]
Robert J Aumann, Michael Maschler, and Richard E Stearns.Repeated games with incomplete information. MIT press, 1995
work page 1995
-
[7]
The evolution of cooperation.science, 211(4489):1390–1396, 1981
Robert Axelrod and William D Hamilton. The evolution of cooperation.science, 211(4489):1390–1396, 1981
work page 1981
-
[8]
Equal pay for all pris- oners.The American mathematical monthly, 104(4):303–305, 1997
Maarten C Boerlijst, Martin A Nowak, and Karl Sigmund. Equal pay for all pris- oners.The American mathematical monthly, 104(4):303–305, 1997
work page 1997
Show all 52 references
-
[9]
The logic of contrition
Maarten C Boerlijst, Martin A Nowak, and Karl Sigmund. The logic of contrition. Journal of Theoretical Biology, 185(3):281–293, 1997. 17
1997
-
[10]
Outlearning extortioners: unbending strategies can foster reciprocal fairness and cooperation.PNAS nexus, 2(6):pgad176, 2023
Xingru Chen and Feng Fu. Outlearning extortioners: unbending strategies can foster reciprocal fairness and cooperation.PNAS nexus, 2(6):pgad176, 2023
2023
-
[11]
Design of zero-determinant strategies and its application to networked repeated games.Science China Information Sciences, 67(10):202201, 2024
Daizhan Cheng and Changxi Li. Design of zero-determinant strategies and its application to networked repeated games.Science China Information Sciences, 67(10):202201, 2024
2024
-
[12]
Zero-determinant strategy in stochastic stackelberg asymmetric security game.Scientific Reports, 13(1):11308, 2023
Zhaoyang Cheng, Guanpu Chen, and Yiguang Hong. Zero-determinant strategy in stochastic stackelberg asymmetric security game.Scientific Reports, 13(1):11308, 2023
2023
-
[13]
Routledge, 2017
Charles Horton Cooley.Human nature and the social order. Routledge, 2017
2017
-
[14]
Ethical implications of ai and robotics in healthcare: A review
Chukwuka Elendu, Dependable C Amaechi, Tochi C Elendu, Klein A Jingwa, Osi- nachi K Okoye, Minichimso John Okah, John A Ladele, Abdirahman H Farah, and Hameed A Alimi. Ethical implications of ai and robotics in healthcare: A review. Medicine, 102(50):e36671, 2023
2023
-
[15]
Emotions and economic theory.Journal of economic literature, 36(1):47– 74, 1998
Jon Elster. Emotions and economic theory.Journal of economic literature, 36(1):47– 74, 1998
1998
-
[16]
Introduction to reinforcement learning.Feuer- riegel, S., Hartmann, J., Janiesch, C., and Zschech, P.(2024)
Damien Ernst and Arthur Louette. Introduction to reinforcement learning.Feuer- riegel, S., Hartmann, J., Janiesch, C., and Zschech, P.(2024). Generative ai. Busi- ness and Information Systems Engineering, 66(1):111–126, 2024
2024
-
[17]
A game-theoretic control approach to mitigate cyber switching attacks in smart grid systems
Abdallah K Farraj, Eman M Hammad, Ashraf Al Daoud, and Deepa Kundur. A game-theoretic control approach to mitigate cyber switching attacks in smart grid systems. In2014 IEEE International Conference on Smart Grid Communications (SmartGridComm), pages 958–963. IEEE, 2014
2014
-
[18]
Cognitive dissonance.Scientific American, 207(4):93–106, 1962
Leon Festinger. Cognitive dissonance.Scientific American, 207(4):93–106, 1962
1962
-
[19]
Influence of opinion dynamics on the evo- lution of games.PloS one, 7(11):e48916, 2012
Floriana Gargiulo and Jos´ e J Ramasco. Influence of opinion dynamics on the evo- lution of games.PloS one, 7(11):e48916, 2012
2012
-
[20]
Psychological games and sequential rationality.Games and economic Behavior, 1(1):60–79, 1989
John Geanakoplos, David Pearce, and Ennio Stacchetti. Psychological games and sequential rationality.Games and economic Behavior, 1(1):60–79, 1989
1989
-
[21]
Zero-determinant strategies in finitely repeated n- player games.IFAC-PapersOnLine, 52(3):150–155, 2019
Alain Govaert and Ming Cao. Zero-determinant strategies in finitely repeated n- player games.IFAC-PapersOnLine, 52(3):150–155, 2019
2019
-
[22]
Zero-determinant strategies in repeated multiplayer social dilemmas with discounted payoffs.IEEE Transactions on Automatic Control, 66(10):4575–4588, 2020
Alain Govaert and Ming Cao. Zero-determinant strategies in repeated multiplayer social dilemmas with discounted payoffs.IEEE Transactions on Automatic Control, 66(10):4575–4588, 2020
2020
-
[23]
Payoff control in the iterated prisoner’s dilemma
Dong Hao, Kai Li, and Tao Zhou. Payoff control in the iterated prisoner’s dilemma. arXiv preprint arXiv:1807.06666, 2018. 18
2018 arXiv
-
[24]
Extortion under uncertainty: Zero- determinant strategies in noisy games.Physical Review E, 91(5):052803, 2015
Dong Hao, Zhihai Rong, and Tao Zhou. Extortion under uncertainty: Zero- determinant strategies in noisy games.Physical Review E, 91(5):052803, 2015
2015
-
[25]
Memory-n strategies of direct reciprocity.Proceedings of the National Academy of Sciences, 114(18):4715–4720, 2017
Christian Hilbe, Luis A Martinez-Vaquero, Krishnendu Chatterjee, and Martin A Nowak. Memory-n strategies of direct reciprocity.Proceedings of the National Academy of Sciences, 114(18):4715–4720, 2017
2017
-
[26]
PhD thesis, The George Washington University, 2019
Qin Hu.Enhancing crowdsourcing with the zero-determinant game theory. PhD thesis, The George Washington University, 2019
2019
-
[27]
Reinforcement learning: A survey.Journal of artificial intelligence research, 4:237–285, 1996
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore. Reinforcement learning: A survey.Journal of artificial intelligence research, 4:237–285, 1996
1996
-
[28]
Academic press, 2014
Samuel Karlin.A first course in stochastic processes. Academic press, 2014
2014
-
[29]
Elsevier, 1981
Samuel Karlin and Howard E Taylor.A second course in stochastic processes. Elsevier, 1981
1981
-
[30]
Environmental quality and population welfare in markovian eco-evolutionary dynamics.Applied Mathematics and Computation, 431:127309, 2022
Fanglin Liu and Bin Wu. Environmental quality and population welfare in markovian eco-evolutionary dynamics.Applied Mathematics and Computation, 431:127309, 2022
2022
-
[31]
Coevolution of vaccination behavior and perceived vacci- nation risk can lead to a stag-hunt-like game.Physical Review E, 106(3):034308, 2022
Yuan Liu and Bin Wu. Coevolution of vaccination behavior and perceived vacci- nation risk can lead to a stag-hunt-like game.Physical Review E, 106(3):034308, 2022
2022
-
[32]
Oxford university press, 2006
George J Mailath and Larry Samuelson.Repeated games and reputations: long-run relationships. Oxford university press, 2006
2006
-
[33]
Zero-determinant strategies under observation errors in repeated games.Physical Review E, 102(3):032115, 2020
Azumi Mamiya and Genki Ichinose. Zero-determinant strategies under observation errors in repeated games.Physical Review E, 102(3):032115, 2020
2020
-
[34]
Continuous opinions and discrete actions in opinion dynamics problems.International Journal of Modern Physics C, 19(04):617–624, 2008
Andr´ e CR Martins. Continuous opinions and discrete actions in opinion dynamics problems.International Journal of Modern Physics C, 19(04):617–624, 2008
2008
-
[35]
A strategy of win-stay, lose-shift that outperforms tit-for-tat in the prisoner’s dilemma game.Nature, 364(6432):56–58, 1993
Martin Nowak and Karl Sigmund. A strategy of win-stay, lose-shift that outperforms tit-for-tat in the prisoner’s dilemma game.Nature, 364(6432):56–58, 1993
1993
-
[36]
Five rules for the evolution of cooperation.science, 314(5805):1560–1563, 2006
Martin A Nowak. Five rules for the evolution of cooperation.science, 314(5805):1560–1563, 2006
2006
-
[37]
Evolutionary games and spatial chaos.nature, 359(6398):826–829, 1992
Martin A Nowak and Robert M May. Evolutionary games and spatial chaos.nature, 359(6398):826–829, 1992
1992
-
[38]
A simple rule for the evolution of cooperation on graphs and social networks.Nature, 441(7092):502–505, 2006
Hisashi Ohtsuki, Christoph Hauert, Erez Lieberman, and Martin A Nowak. A simple rule for the evolution of cooperation on graphs and social networks.Nature, 441(7092):502–505, 2006
2006
-
[39]
Direct reciprocity on graphs.Journal of theoretical biology, 247(3):462–470, 2007
Hisashi Ohtsuki and Martin A Nowak. Direct reciprocity on graphs.Journal of theoretical biology, 247(3):462–470, 2007. 19
2007
-
[40]
Zero-determinant strategies in iterated public goods game.Scientific reports, 5(1):13096, 2015
Liming Pan, Dong Hao, Zhihai Rong, and Tao Zhou. Zero-determinant strategies in iterated public goods game.Scientific reports, 5(1):13096, 2015
2015
-
[41]
Iterated prisoner’s dilemma contains strate- gies that dominate any evolutionary opponent.Proceedings of the National Academy of Sciences, 109(26):10409–10413, 2012
William H Press and Freeman J Dyson. Iterated prisoner’s dilemma contains strate- gies that dominate any evolutionary opponent.Proceedings of the National Academy of Sciences, 109(26):10409–10413, 2012
2012
-
[42]
Emotion, cognition, and decision making.Cognition & emotion, 14(4):433–440, 2000
Norbert Schwarz. Emotion, cognition, and decision making.Cognition & emotion, 14(4):433–440, 2000
2000
-
[43]
Payoff control in multichannel games: Influencing opponent learning evolution
Juan Shi, Chen Chu, Guoxi Fan, Die Hu, Jinzhuo Liu, Zhen Wang, and Shuyue Hu. Payoff control in multichannel games: Influencing opponent learning evolution. IEEE Transactions on Cybernetics, 2025
2025
-
[44]
Evolutionary dynam- ics with game transitions.Proceedings of the National Academy of Sciences, 116(51):25398–25404, 2019
Qi Su, Alex McAvoy, Long Wang, and Martin A Nowak. Evolutionary dynam- ics with game transitions.Proceedings of the National Academy of Sciences, 116(51):25398–25404, 2019
2019
-
[45]
Competition of tolerant strategies in the spatial public goods game.New Journal of Physics, 18(8):083021, 2016
Attila Szolnoki and Matjaˇ z Perc. Competition of tolerant strategies in the spatial public goods game.New Journal of Physics, 18(8):083021, 2016
2016
-
[46]
An incentive mechanism for federated learning: A continuous zero-determinant strategy approach.IEEE/CAA Journal of Automatica Sinica, 11(1):88–102, 2024
Changbing Tang, Baosen Yang, Xiaodong Xie, Guanrong Chen, Mohammed AA Al- Qaness, and Yang Liu. An incentive mechanism for federated learning: A continuous zero-determinant strategy approach.IEEE/CAA Journal of Automatica Sinica, 11(1):88–102, 2024
2024
-
[47]
Individual costs and societal benefits of interventions during the covid-19 pandemic.Proceedings of the National Academy of Sciences, 120(24):e2303546120, 2023
Arne Traulsen, Simon A Levin, and Chadi M Saad-Roy. Individual costs and societal benefits of interventions during the covid-19 pandemic.Proceedings of the National Academy of Sciences, 120(24):e2303546120, 2023
2023
-
[48]
Memory-two zero-determinant strategies in repeated games.Royal Society open science, 8(5):202186, 2021
Masahiko Ueda. Memory-two zero-determinant strategies in repeated games.Royal Society open science, 8(5):202186, 2021
2021
-
[49]
Necessary and sufficient condition for the existence of zero- determinant strategies in repeated games.Journal of the Physical Society of Japan, 91(8):084801, 2022
Masahiko Ueda. Necessary and sufficient condition for the existence of zero- determinant strategies in repeated games.Journal of the Physical Society of Japan, 91(8):084801, 2022
2022
-
[50]
An oscillating tragedy of the commons in replicator dynamics with game-environment feedback.Proceedings of the National Academy of Sciences, 113(47):E7518–E7525, 2016
Joshua S Weitz, Ceyhun Eksin, Keith Paarporn, Sam P Brown, and William C Ratcliff. An oscillating tragedy of the commons in replicator dynamics with game-environment feedback.Proceedings of the National Academy of Sciences, 113(47):E7518–E7525, 2016
2016
-
[51]
Learning multimodal confidence for intention recognition in human- robot interaction.IEEE Robotics and Automation Letters, 2024
Xiyuan Zhao, Huijun Li, Tianyuan Miao, Xianyi Zhu, Zhikai Wei, Lifen Tan, and Aiguo Song. Learning multimodal confidence for intention recognition in human- robot interaction.IEEE Robotics and Automation Letters, 2024
2024
-
[52]
A two-layer model for coevolving opinion dynamics and collective decision-making in complex social systems.Chaos: An Interdisciplinary Journal of Nonlinear Science, 30(8), 2020
Lorenzo Zino, Mengbin Ye, and Ming Cao. A two-layer model for coevolving opinion dynamics and collective decision-making in complex social systems.Chaos: An Interdisciplinary Journal of Nonlinear Science, 30(8), 2020. 20
2020
Reviewed July 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.