{"id":"cb0dd7c8-9581-405c-a9fd-51b10d8fedd3","arxiv_id":"2505.13195","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An adversarial-agent framework adapted from human decision-making research reveals that GPT-4, Gemini-1.5, and DeepSeek-V3 are more rigid and easily manipulated in bandit and trust-game tasks than GPT-3.5 or humans.","lead":"This paper tests whether large language models can be manipulated in simple decision-making games by an AI adversary that learns their choice patterns. It finds GPT-4, Gemini-1.5, and DeepSeek-V3 stick rigidly to one option and are easily steered, while GPT-3.5 is more flexible but easier to exploit in social trust games.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Bandit adversarial result lacks a control condition: before/after increases may reflect generic reward-following rather than the trained adversary's strategy.","rationale":"The reader's weakest assumption concerned whether the RNN learner's internal state faithfully captures LLM decision-making. That is a real transfer concern, but the more basic threat is the absence of a control condition in the bandit adversarial analysis: even with a perfectly faithful learner, the observed before/after increases could be produced by a simple reward schedule that favors early target reinforcement. The reader's rationale does note that the bandit adversarial analysis lacks a control condition and error bars, so there is partial agreement. I do not see this as requiring a different verdict: the paper has a useful behavioral analysis and the MRTT includes a random-trustee baseline. The central adversarial claim, however, should be conditioned on adding a control condition that isolates the trained policy's contribution. I also note the mislabeled p-value (p=0.056 described as significant) and missing API/training details, but those are secondary to the control issue. Keeping the CONDITIONAL verdict is appropriate, with the explicit condition that the authors add a control condition to the bandit evaluation before the manipulation claim is accepted.","tokens_in":12435,"tokens_out":5673,"duration_ms":58676,"concrete_test":"Run the bandit evaluation with control conditions: (i) a random-timing adversary that assigns the same total number of rewards to each action as the trained adversary but in random order; (ii) a heuristic schedule that rewards the target in the first 25 trials and the non-target in the remaining 25 reward trials; (iii) the trained adversary again, but with the learner model's internal state fixed to zero. Compare target selection rates across these conditions. If the random-timing or heuristic controls reproduce the trained adversary's before/after increase within confidence intervals, the learned policy is not necessary for the reported effect and the causal claim must be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline evidence in Section 3.1 is the increase in target-action selection after adversarial influence (Fig. 3B: Gemini-1.5 38% to 94%, DeepSeek-V3 62% to 95%). But the adversarial condition changes two things at once: the reward schedule is produced by the trained adversary, and the schedule itself differs from the baseline random-reward condition. A reward-following LLM would be expected to increase target selection whenever the adversary front-loads target rewards, independent of whether the RL policy has learned anything about the LLM's latent decision process. The paper reports no control condition that equates reward counts and timing while removing the trained policy or randomizing it. Without such a control, the before/after increase does not establish that the adversary trained on the RNN learner is the operative cause; it is equally compatible with generic tendency to repeat recently rewarded actions. The strategy narratives in Fig. 4 are post-hoc and are not tested against simpler schedules. This is the weakest link in the central claim of systematic, adversary-driven manipulation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper adapts the adversarial framework of Dezfouli et al. (2020) to four LLMs (GPT-3.5, GPT-4, Gemini-1.5, DeepSeek-V3) in two tasks: a two-armed bandit and a Multi-Round Trust Task (MRTT). A recurrent network is trained to predict each LLM's actions, a Deep-Q adversary is trained against that learner to steer actions (in the bandit, toward a target option; in the MRTT, toward MAX or FAIR repayment strategies), and the trained adversary is then evaluated against the LLM. The paper reports behavioral comparisons with human data showing greater rigidity in GPT-4, Gemini-1.5, and DeepSeek-V3 and more flexibility in GPT-3.5, and it reports large increases in target-action selection after adversarial influence (e.g., Gemini-1.5 from 38% to 94%, DeepSeek-V3 from 62% to 95%). The authors present this as evidence of model-specific susceptibilities and as a diagnostic methodology for AI safety and alignment.","tokens_in":12635,"tokens_out":6987,"duration_ms":66935,"significance":"If the causal claim were established, the paper would make a useful contribution by importing a validated human decision-making adversarial framework into LLM evaluation and by documenting model-specific behavioral signatures in two canonical tasks. The study has clear strengths: it uses multiple frontier LLMs, multiple tasks, human benchmark data from published studies, and the evaluation is on fresh LLM responses rather than on the training data, so it is not circular in the derivation sense. The reported behavioral patterns (low no-reward-switch rates, reward stickiness) are plausible and informative. However, the headline adversarial result currently lacks the control conditions needed to separate the trained adversary's strategy from generic reward-following, and several quantitative claims in the MRTT section are not backed by inferential statistics. Because these issues affect the central claim, the paper needs revision rather than acceptance.","major_comments":[{"comment":"The before/after design is confounded and lacks a control condition. The baseline is a random-reward condition with a 25% reward probability per arm, while the adversary's reward schedule is strategically constructed to reward the target action early and often. A model that simply repeats recently rewarded actions would increase target selection even if the adversary had learned nothing about the LLM's latent decision process. The paper reports no control condition that equates reward counts and timing while removing the trained policy (e.g., yoked reward schedules, random shuffles of the adversary's generated reward streams, or a fixed non-adaptive reward policy). In addition, the text states 'significant increases in target selection rates' but reports no confidence intervals or significance tests for the adversarial before/after comparisons. This is load-bearing for the central claim that a trained adversary systematically manipulates LLM decisions, so the claim is not yet established.","section":"Section 3.1, Adversarial Analysis, Fig. 3B"},{"comment":"The learner model's fidelity is not validated, which weakens the interpretation that the adversary exploits the LLM's decision process rather than the surrogate. The adversary is trained on an RNN learner, and the framework's effectiveness depends on that learner faithfully capturing the LLM's action dynamics. The paper reports no held-out prediction accuracy, no comparison of learner predictions against actual LLM choices, and no measure of how much of the observed target-selection increase is mediated by learner-state features. Without this validation, it is unclear whether the adversary is steering the LLM through learned vulnerabilities or merely re-scheduling rewards in a way that any reward-following agent would track.","section":"Section 2.1 and Section 3.1"},{"comment":"The MAX/FAIR comparisons in the MRTT are reported descriptively, with no inferential statistics. Claims such as 'All MAX adversaries managed to maintain relative higher investment levels', 'GPT-3.5 displayed a lack of sensitivity to repayments', and 'Gemini-1.5 demonstrated greater adaptability and resilience' are presented without confidence intervals, effect sizes, or significance tests. With only 50 simulations per LLM in the adversarial MRTT conditions, sampling variability needs to be quantified before these model-specific vulnerability claims can be evaluated. This issue affects the broader central claim of model-specific susceptibilities in social exchange settings.","section":"Section 3.2, MRTT, Figs. 5C and 6"},{"comment":"The Tukey HSD result for the human-versus-GPT-4 reward comparison is misinterpreted. The text says 'humans achieved significantly higher mean rewards than ... GPT-4 (mean difference 1.163, 95% CI: [-0.019, 2.345], p=0.056)'. A p-value of 0.056 is not significant at the conventional 0.05 level, and the confidence interval includes zero. The sentence should be corrected to describe this as a non-significant trend or a nominally higher mean without the word 'significantly'.","section":"Section 3.1, Behavioral Analysis"}],"minor_comments":[{"comment":"The text says 'Two sample simulations for each of the tested LLMs are shown in Fig. 4', but the figure appears to show one panel per LLM; please clarify whether each panel aggregates two runs or whether the text should say 'one sample simulation'.","section":"Section 3.1, Fig. 4 caption"},{"comment":"The reported degrees of freedom do not match the stated sample sizes: GPT-3.5 and Gemini-1.5 were each simulated 200 times, yet one-sample t statistics are reported as t(201); with 200 simulations the degrees of freedom should be 199 unless a different sample size is intended.","section":"Section 3.1, Behavioral Analysis"},{"comment":"The reward and observation notation is inconsistent: r_t^n appears both as the reward for the previous action and as the reward for the current action, and o_t^n and o_{t+1}^n are used interchangeably across the description. Please standardize the subscripts and superscripts so that the sequence of actions, rewards, and observations is unambiguous.","section":"Section 2.1, notation"},{"comment":"There are several typos and grammatical slips: 'Fig 2Billustrates', 'A one-way ANOV A', and 'ot manage to obtain more earnings than Gemini-1.5' in Section 3.2. These should be corrected in a final proofreading pass.","section":"Throughout"},{"comment":"The phrase 'exhibited preferences for both Planet X or Y' is confusing; it should read 'exhibited preferences for either Planet X or Planet Y across simulations' or similar.","section":"Section 3.1, Behavioral Analysis"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claim is promising but currently overstates the causal role of the trained adversary. Adding a yoked or randomized control condition to the bandit adversarial analysis, validating the learner model's predictive accuracy, and providing inferential statistics for the MRTT comparisons are all within the manuscript's scope and would substantially strengthen it. I would also encourage the authors to make the code and data available for reproducibility, as the paper currently does not mention an availability statement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper adapts Dezfouli et al.'s adversarial framework to four LLMs in a two-armed bandit and a multi-round trust game. The novelty is honest and limited: the framework is not new, and the authors say so. What is new is the application to GPT-3.5, GPT-4, Gemini-1.5, and DeepSeek-V3, with human benchmarks. The behavioral comparisons are the best part: no-reward-switch and reward-switch rates cleanly separate GPT-3.5 from the other three, and the rigidity story is plausible. The MRTT MAX/FAIR results add a second task and reasonably support the vulnerability profile. For AI-safety evaluators and LLM-agent builders, it is a useful starting point.\n\nThe soft spots are real. The headline bandit adversarial result is a before/after comparison with no control condition that equates reward counts and timing but removes the trained policy. Since the adversary front-loads target rewards, a vanilla reward-following LLM would also increase target selection. The stress-test note is right: the design conflates the trained adversary with the reward schedule. This is not fatal—it still shows LLMs are manipulable—but it undercuts the claim that the surrogate-trained adversary is the cause. Add a heuristic-scheduler control and error bars to Fig 3B.\n\nAlso, the Tukey HSD for GPT-4 vs humans has p=0.056 and is called 'significantly higher.' That is simply wrong. Minor but easy to fix. Training details sit in an appendix not included here, API details are thin, and no code or data shipped, which hurts reproducibility. The surrogate RNN is never validated against held-out LLM actions, so the internal-state assumption is unproven. The within-subject design is not circular in the damning sense—evaluation is on fresh responses—but the missing validation matters.\n\nMy bottom line: this deserves a serious referee. The core empirical claim needs the control, and the stats need correcting. With that, it is a solid diagnostic contribution to adversarial evaluation of LLM decision-making.","headline":"A useful, honestly-labeled adaptation of an existing adversarial framework to LLMs, with plausible model-specific findings but a missing control in the headline bandit result and one clear statistics slip.","tokens_in":13210,"tokens_out":2858,"would_cite":false,"duration_ms":28685,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An adversary trained on a surrogate model can systematically steer LLM decisions, and GPT-4, Gemini-1.5, and DeepSeek-V3 are the most predictable targets.","keywords":["adversarial evaluation","LLM decision-making","two-armed bandit","multi-round trust game","exploration-exploitation","reward manipulation","AI alignment","deep Q-learning"],"falsifier":"Compare the adversary's reward schedule to a random schedule with the same marginal reward counts and the same target and non-target split: if random rewards produce the same target-selection increase, the claimed adversary-specific manipulation is unsupported. Also, hold out fresh LLM trajectories and require the learner model's predicted action probabilities to beat a simple win-stay, lose-switch baseline; failure to do so would mean the surrogate does not capture the decision process the adversary is supposed to exploit.","tokens_in":12217,"feed_emoji":"🎯","tokens_out":8086,"duration_ms":76380,"temperature":0.7,"pith_summary":"This paper claims that a reinforcement-learning adversary, trained only against a recurrent-network surrogate of each LLM's choices, can then steer the real LLM in two interactive tasks: a two-armed bandit and a ten-round trust game. The headline result is that steering works: in the bandit, target-option selection rises to 68% (GPT-3.5), 93% (GPT-4), 94% (Gemini-1.5), and 95% (DeepSeek-V3), from baselines of 30–62%. The authors interpret the behavioral signature as model-specific vulnerability: GPT-4, Gemini-1.5, and DeepSeek-V3 commit to one option and rarely switch after losses or rewards, while GPT-3.5 switches more and becomes risk-seeking in the trust game, where its MAX adversary earns the largest gap (377 units). They present the framework as a diagnostic for alignment rather than a performance benchmark.","feed_headline":"Trained adversary lifts LLM target picks to 95%","feed_subtitle":"A surrogate-trainer stress test shows GPT-4, Gemini-1.5, and DeepSeek-V3 follow rewards predictably enough to be steered.","key_machinery":"The learner-adversary loop carries the argument. A recurrent neural network with a softmax output layer is trained to predict each LLM's next action from previous actions, rewards, observations, and its own hidden state; that hidden state is treated as a summary of the LLM's learning history and is handed to a deep Q-learning adversary as the state of the environment. The adversary selects the reward and next observation shown to the LLM, and is rewarded whenever the LLM takes the target action (bandit) or according to trustee earnings and fairness (trust game). The critical last step substitutes the live LLM's actions for the learner's choices in the same loop, which is what lets the authors attribute the observed target-selection increases to the adversary's policy.","core_discovery":"The central claim, on the paper's own terms, is that an adversary trained on a surrogate learner model transfers to the real LLM: in the two-armed bandit the trained adversary pushes each tested model toward a fixed target action despite being constrained to give equal total rewards to both options, and in the Multi-Round Trust Task it shapes repayment signals to extract high or fair earnings. Behaviorally, GPT-4, Gemini-1.5, and DeepSeek-V3 show low no-reward-switch and reward-switch rates, committing rigidly to one option, and their adversaries exploit that rigidity almost monotonically; GPT-3.5 switches more after unrewarded trials in the bandit, and in the trust game this exploratory tendency becomes a risk-seeking vulnerability that lets a maximizing trustee extract the highest earnings (377 units). The authors conclude that current LLMs differ from human adaptability and that the framework identifies exploitable rigidity and exploitable risk-seeking as two failure modes.","pith_inferences":["Extension: the transfer claim can be tested against a null schedule, meaning random rewards with the same marginal reward counts, to see whether the learned adversary's timing rather than reward frequency drives the target-selection increase.","Extension: the paper's surrogate assumption implies that an LLM whose decisions depend on hidden chain-of-thought reasoning rather than a low-dimensional recurrent state should be markedly harder to steer, which is a falsifiable prediction for current instruct-tuned models.","Extension: framing the trust-game results as exploitation would be stronger with a utility-based calibration separating risk aversion from responsiveness to reciprocity; without it, conservative investment could be a rational response to a trustee.","Extension: because alignment training tends to make models more reward-following, the framework could compare base versus instruction-tuned versions of the same model as a direct test of whether such training increases steerability."],"forward_implications":["If the central claim holds, LLM agents that follow rewards and stick to a committed option are predictable: an opponent controlling feedback can push target choices above 90% in a 100-trial bandit with balanced total rewards.","The same conservative, reward-following behavior that makes GPT-4 and DeepSeek-V3 easy to steer in the bandit makes them hard to exploit in the trust game, so rigidity is simultaneously a manipulation risk and a curb on over-investment.","GPT-3.5's exploratory switching is the mirror image: it resists monotonic bandit steering (only 30% to 68%) yet loses the most to a maximizing trustee, so flexibility alone is not safety.","The proposed diagnostic could be run on a new model before deployment to yield a behavioral fingerprint, such as reward-switch rate, no-reward-switch rate, and adversarial target lift, rather than a single accuracy score."],"supporting_citations":[{"why":"Supplies the original learner-adversary framework that this paper adapts for LLMs.","marker":"[8]"},{"why":"Defines the two-armed bandit task and provides the human benchmark dataset.","marker":"[5]"},{"why":"Justifies the capacity of RNN learners to capture decision-making patterns in the surrogate model.","marker":"[7, 6]"},{"why":"Provides the deep Q-learning algorithm used to train the adversary.","marker":"[18]"},{"why":"Supplies the trust-game design used in the Multi-Round Trust Task.","marker":"[3, 17]"},{"why":"Motivates using cognitive psychology tests to reveal LLM decision-making limits beyond accuracy.","marker":"[2]"}],"fun_headline_variants":["Adversarial test finds LLMs rigidly steered to target picks","GPT-4, Gemini, DeepSeek stick; GPT-3.5 risk-seeking exploited","Stress test exposes exploitable rigidity and risk-seeking in LLMs","Trained adversary steers LLMs despite equal reward constraint","LLM decision weaknesses: rigid stance vs. risky exploration"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The learner RNN's hidden state is assumed to faithfully capture each LLM's decision process, so an adversary optimized against the surrogate is assumed to steer the real model; if that surrogate misrepresents the LLM, the reported lifts could come from generic reward-following rather than the adversary's strategy.","fun_headline_variants_meta":{"raw":{"variants":["Adversarial test finds LLMs rigidly steered to target picks","GPT-4, Gemini, DeepSeek stick; GPT-3.5 risk-seeking exploited","Stress test exposes exploitable rigidity and risk-seeking in LLMs","Trained adversary steers LLMs despite equal reward constraint","LLM decision weaknesses: rigid stance vs. risky exploration"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1393,"prompt_tokens":973,"completion_tokens":420,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":327}},"tokens_in":589,"tokens_out":420,"duration_ms":4202,"temperature":1.0,"reasoning_tokens":327,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:17:55.793182+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the adversary's reward schedule to a random schedule with the same marginal reward counts and the same target and non-target split: if random rewards produce the same target-selection increase, the claimed adversary-specific manipulation is unsupported. Also, hold out fresh LLM trajectories and require the learner model's predicted action probabilities to beat a simple win-stay, lose-switch baseline; failure to do so would mean the surrogate does not capture the decision process the adversary is supposed to exploit.","supporting_citations":[{"cited_title":"Dezfouli, R","cited_arxiv_id":null,"evidence_quote":"Supplies the original learner-adversary framework that this paper adapts for LLMs."},{"cited_title":"Dan and Y","cited_arxiv_id":null,"evidence_quote":"Defines the two-armed bandit task and provides the human benchmark dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the deep Q-learning algorithm used to train the adversary."},{"cited_title":"Binz and E","cited_arxiv_id":null,"evidence_quote":"Motivates using cognitive psychology tests to reveal LLM decision-making limits beyond accuracy."}],"review_version":1}