{"id":"f4f8974a-3b29-48ed-9338-2a33e93a5fd2","arxiv_id":"2607.00625","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"Positive and negative determinant strategies enable unilateral payoff control in repeated games with behavior-value inconsistency costs, where zero-determinant strategies cease to exist.","lead":"This paper introduces positive and negative determinant strategies in repeated games where agents incur internal costs when their actions mismatch their internal values. These strategies enable unilateral control over payoffs, such as capping an opponent's payoff or ensuring a higher payoff than the opponent, outperforming zero-determinant strategies under inconsistency costs.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader already flags the unshown proofs and the cost-structure assumption as the key uncertainty; no additional technical gap is visible in the abstract-level description of the argument.","tokens_in":1823,"tokens_out":174,"duration_ms":31781,"concrete_test":"Extract the exact cost function and modified payoff expressions from the full manuscript; substitute into the classic ZD linear-condition equation and verify whether any solution remains for generic cost parameters.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on a specific internal-cost model for behavior-value inconsistency being sufficient to rule out ZD strategies while permitting the new determinant strategies. No internal inconsistency, hidden assumption in the payoff construction, or non-general step in the claimed non-existence proof can be isolated from the provided description.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces a repeated-game framework incorporating an internal cost for behavior-value inconsistency. It claims to prove that zero-determinant (ZD) strategies cease to exist under this cost model. In their place, the authors define a new class of 'positive/negative determinant strategies' that enforce an affine combination of the two players' average payoffs to lie strictly above or below zero. These strategies are asserted to permit unilateral control of the opponent's payoff (negative determinant) or to guarantee a higher payoff than the opponent (positive determinant), with superior control ability relative to classic ZD strategies.","tokens_in":1880,"tokens_out":422,"duration_ms":16994,"significance":"If the claimed non-existence result and the construction of the new determinant strategies hold, the work would usefully extend the theory of unilateral payoff control by incorporating an internal consistency cost that is relevant to models of intelligent agents. The paper would thereby highlight a limitation of standard ZD analyses that omit such costs. No machine-checked proofs, reproducible code, or parameter-free derivations are described.","major_comments":[{"comment":"The abstract and reader's summary assert proofs of ZD non-existence and of the control properties of positive/negative determinant strategies, yet no derivations, payoff matrices, recurrence relations, or explicit strategy definitions appear in the provided text. Without these, the central claims cannot be verified and the soundness assessment remains low.","section":null},{"comment":"The framework defines the new strategies relative to a specific inconsistency-cost model. It is unclear whether the claimed affine-control property follows independently of the particular functional form chosen for the cost or whether it reduces to a fitted parameter (cf. the reader's note on free_parameters = ['inconsistency cost']).","section":null}],"minor_comments":[{"comment":"The abstract contains minor grammatical issues ('provided what agents do differ', 'termed as') that should be corrected for clarity.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed report and the opportunity to clarify our contributions. We address the two major comments point by point below. The full manuscript contains the requested derivations, but we will revise to make them more prominent and to address generality concerns.","responses":[{"response":"The full manuscript (Sections 3 and 4) contains the non-existence proof for ZD strategies (Theorem 1), explicit strategy definitions for positive/negative determinant strategies, payoff matrices for the iterated Prisoner's Dilemma, and the recurrence relations used to derive the affine control property. We will revise the submission to include a dedicated appendix with all derivations, matrices, and strategy update rules to facilitate verification.","revision_made":"yes","referee_comment":"The abstract and reader's summary assert proofs of ZD non-existence and of the control properties of positive/negative determinant strategies, yet no derivations, payoff matrices, recurrence relations, or explicit strategy definitions appear in the provided text. Without these, the central claims cannot be verified and the soundness assessment remains low."},{"response":"The non-existence of ZD strategies holds for any strictly positive inconsistency cost (i.e., whenever behavior differs from internal value). The affine-control property of positive/negative determinant strategies is derived under this general condition and does not depend on a specific functional form; the linear cost used in the paper is for concrete illustration only. The inconsistency cost is an explicit model parameter, not a fitted one. We will add a new subsection clarifying the generality of the results and the role of the cost parameter.","revision_made":"partial","referee_comment":"The framework defines the new strategies relative to a specific inconsistency-cost model. It is unclear whether the claimed affine-control property follows independently of the particular functional form chosen for the cost or whether it reduces to a fitted parameter (cf. the reader's note on free_parameters = ['inconsistency cost'])."}],"tokens_in":1369,"tokens_out":414,"duration_ms":19902,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that introducing an internal cost when actions mismatch an agent's internal value removes zero-determinant strategies and replaces them with a new class called positive and negative determinant strategies. These let one player force an affine combination of average payoffs above or below zero, which the abstract says enables unilateral control such as capping the opponent's payoff or securing a higher payoff than the opponent.\n\nThe work does extend the ZD framework by incorporating this inconsistency cost, which is a plausible addition for modeling agents that care about internal consistency. It positions the new strategies as emerging only under this cost and claims they deliver stronger control than classic ZD ones. That is a clear, focused move beyond the prior literature.\n\nThe soft spot is that the abstract asserts proofs of ZD non-existence and the enforcement properties without showing equations, steps, or verification. Without those derivations it is hard to judge whether the results are independent of the cost parameter or whether the non-existence claim holds generally. The cost itself appears as a modeling choice that drives the whole distinction.\n\nThis is for people already working on repeated games and ZD strategies who want to add internal state costs. A reader familiar with Press-Dyson style work would see the connection immediately.\n\nIt should go to peer review so the derivations can be examined directly.","headline":"The paper adds an internal cost for behavior-value inconsistency in repeated games, claims this eliminates ZD strategies, and introduces positive/negative determinant strategies that enforce affine payoff combinations instead.","tokens_in":2384,"tokens_out":321,"would_cite":false,"duration_ms":19762,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Behavior-value inconsistency eliminates zero-determinant strategies but enables positive and negative determinant strategies for unilateral payoff control.","keywords":["repeated games","zero-determinant strategies","behavior-value inconsistency","positive determinant strategy","negative determinant strategy","unilateral payoff control","direct reciprocity"],"falsifier":"A calculation or simulation showing that a zero-determinant strategy still succeeds in enforcing its payoff relation, or that a positive/negative determinant strategy fails to enforce the claimed affine combination, once the inconsistency cost is subtracted from the payoffs.","tokens_in":2695,"feed_emoji":"","tokens_out":707,"duration_ms":27917,"temperature":0.7,"pith_summary":"The paper builds a repeated-game model in which each agent pays an internal cost whenever its chosen action differs from its internal thought or value. Under this cost, the authors demonstrate that classic zero-determinant strategies cannot exist. They instead identify a new family of strategies, called positive and negative determinant strategies, that force an affine combination of the two players' long-run average payoffs to lie strictly above or below zero. These strategies let one player unilaterally drive the opponent's payoff below any chosen threshold or ensure its own payoff exceeds the opponent's. The framework matters because it shows how internal consistency costs reshape the possibilities for one-sided influence in repeated interactions.","feed_headline":"Inconsistency cost removes ZD strategies from repeated games","feed_subtitle":"Positive and negative determinant strategies replace them with stronger unilateral payoff control.","key_machinery":"Positive/negative determinant strategy, which forces an affine combination of the two players' average payoffs to be positive or negative.","core_discovery":"We prove that ZD strategy does not exist if the cost via behavior-value inconsistency is present. Instead, we find a new class of repeated strategies that enforce a unilateral payoff control, which is termed as positive/negative determinant strategy. The found strategy allows an individual to enforce an affine combination of two individuals' average payoffs above/below zero. Consequently, a focal individual is able to unilaterally control the opponent's payoff below a given value via negative determinant strategy, and a focal individual is able to get more payoff than the opponent via positive determinant strategy. We also find that the control ability of positive/negative determinant strate","pith_inferences":["Models of reciprocity that assume no internal costs may systematically overstate the reach of zero-determinant strategies.","Artificial agents could adopt positive/negative determinant strategies to achieve similar unilateral controls while remaining internally consistent.","Empirical tests could check whether human subjects in repeated games behave as if they incur costs for action-value mismatches."],"forward_implications":["Zero-determinant strategies cannot be used once inconsistency costs are included in the payoff structure.","A focal player using a negative determinant strategy can unilaterally force the opponent's average payoff below any chosen level.","A focal player using a positive determinant strategy can unilaterally ensure its own average payoff exceeds the opponent's.","The payoff control achieved by positive/negative determinant strategies is strictly stronger than the control achieved by zero-determinant strategies."],"fun_headline_variants":["Inconsistency costs eliminate ZD strategies","Behavior value mismatch blocks ZD strategies","New determinants enable unilateral payoff control","ZD strategies absent under inconsistency costs","Positive negative determinants enforce payoff combos"],"cache_read_input_tokens":64,"weakest_assumption_plain":"An agent pays an internal cost exactly when its behavior differs from its internal thought or value, and this cost structure is enough to remove ZD strategies while permitting the new determinant strategies.","fun_headline_variants_meta":{"raw":{"variants":["Inconsistency costs eliminate ZD strategies","Behavior value mismatch blocks ZD strategies","New determinants enable unilateral payoff control","ZD strategies absent under inconsistency costs","Positive negative determinants enforce payoff combos"]},"model":"grok-4.3","cost_usd":0.004808,"raw_usage":{"total_tokens":2302,"prompt_tokens":704,"num_sources_used":0,"completion_tokens":46,"cost_in_usd_ticks":48078000,"prompt_tokens_details":{"text_tokens":704,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1552,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":704,"tokens_out":46,"duration_ms":16507,"temperature":1.0,"reasoning_tokens":1552,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-02T04:13:02.232519+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A calculation or simulation showing that a zero-determinant strategy still succeeds in enforcing its payoff relation, or that a positive/negative determinant strategy fails to enforce the claimed affine combination, once the inconsistency cost is subtracted from the payoffs.","supporting_citations":[],"review_version":1}