{"id":"3e95a620-f8be-49bd-8b78-1f3c59117c25","arxiv_id":"2506.06366","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.","lead":"This paper proposes AI Agent Behavioral Science, a framework that studies large language model agents by observing their situated behavior rather than by inspecting model internals. It synthesizes research on individual agents, multi-agent systems, and human-agent interaction, and argues this lens should guide responsible AI.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paradigm assumes stable, reproducible agent behavior, but the paper's own evidence shows extreme prompt and version sensitivity, so its empirical grounding is unestablished.","rationale":"The reader's weakest_assumption correctly identifies the load-bearing premise: the paradigm requires that LLM-based agents exhibit stable, reproducible behavioral regularities that can be studied with human behavioral-science methods. My stress-test confirms this and sharpens it into a falsifiable reliability requirement. The paper's own cited studies supply the threat: prompt-sensitivity findings in Sections 2.1 and 6.2 (e.g., an emoji altering outputs [174]) and order effects in Section 6.3 [156] are exactly the kind of instability that would invalidate behavioral measurement. The paradigm also asserts that behavior emerges from the agentic system, not the model alone; this is only meaningful if the system's contribution is systematic and separable from mere prompt/context variation. Since the paper presents no reliability, generalizability, or scaffold-ablation data, the 'emergent behavior' claim remains a plausible framing rather than an established empirical finding. This does not warrant rejection of a position paper; it warrants the CONDITIONAL verdict already given. The proposed ICC/variance test provides a concrete falsification check: if the same model with the same scaffold yields behavioral measures that are swamped by paraphrase or version noise, the brain-to-model analogy loses its force. I disagree with the reader only on emphasis: the internal Table 2 contradiction, while real, is a fixable editorial issue; the stability question is the deeper threat and should be foregrounded in revisions.","tokens_in":37798,"tokens_out":5866,"duration_ms":57856,"concrete_test":"Compute test-retest reliability of three behavioral measures from the paper (Dictator-Game generosity [103], Prisoner's-Dilemma cooperation [58], belief-updating in game theory [56]) across 20 semantically equivalent prompt paraphrases and 5 snapshots of the same LLM API over a month. If intraclass correlation < 0.5, or if paraphrase/version variance exceeds between-condition variance, then agent behavior is not stable enough to ground the proposed paradigm; if reliability is high, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that AI agents are behavioral entities whose behavior emerges from situated interaction and is not determined by the model alone—requires that agent behavior exhibit stable, reproducible regularities that can be measured, compared, and generalized, as the paper's reliance on human behavioral science implies. Yet the paper's own cited evidence repeatedly shows that LLM outputs are acutely sensitive to arbitrary prompt phrasing, ordering, and model version: Section 2.1 reports context-dependent rationality and 'LLM sensitivity to prompts' [95, 171]; Section 6.2 notes that a single added emoji can significantly alter outputs [174]; Section 6.3 reports order effects in similarity judgments [156]. Behavioral science methods assume test-retest reliability: a 'generous' or 'cooperative' agent must be so reliably across equivalent presentations, not just for one prompt template. The paper never provides such reliability evidence; every summarized result is a single-snapshot sandbox observation. If behavioral measures vary more across paraphrases or model snapshots than across experimental conditions, then what the paradigm calls 'emergent behavior' is largely an artifact of prompt/version variance, and the model-to-behavior analogy loses its empirical grounding. This is an internal-consistency problem, not just a disagreement with consensus: the paper's own selected literature supplies evidence that threatens its core premise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a new research paradigm called \"AI Agent Behavioral Science,\" arguing that LLM-based agents should be studied not only through internal model mechanisms but as behavioral entities whose actions emerge from situated interaction with environments, other agents, and humans. It synthesizes recent empirical work across three settings—individual agents, multi-agent systems, and human-agent interaction—organizes behavioral adaptation methods using a reinterpretation of the Fogg Behavior Model (ability, motivation, trigger), and reframes responsible AI principles (fairness, safety, interpretability, accountability, privacy) as behavioral properties. The paper closes with six proposed research directions. It is a synthesis and position paper rather than a new experimental study.","tokens_in":37962,"tokens_out":5302,"duration_ms":46352,"significance":"If its central premise is accepted, the paper makes a timely contribution by connecting machine behavior, behavioral economics, social simulation, and AI governance under one framework. Its strengths include broad literature coverage, a useful taxonomy of emergent agent behaviors, and a concrete agenda for studying behavioral reliability and adaptation. However, the key conceptual claim that this paradigm is a \"necessary complement\" to model-centric approaches is asserted rather than demonstrated, and the paper's own cited evidence raises substantial concerns about the stability and reproducibility of agent behavior that the paradigm presupposes. The internal contradictions in summarizing key references further weaken the synthesis. With revisions that address these issues, the paper could be a valuable roadmap for an emerging field.","major_comments":[{"comment":"The central claim that AI Agent Behavioral Science is a \"necessary complement\" to model-centric approaches is asserted rather than demonstrated. The paper surveys examples of emergent behavior, but it does not identify a case where a behavioral-level account yields predictions, explanations, or governance insights that are unavailable in principle from model-centric analysis, nor does it state what evidence would count against the necessity claim. The conclusion should be either softened to a \"valuable complement\" or supported by an explicit argument and testable criteria.","section":"Section 1, 8"},{"comment":"The paradigm presupposes that agent behavior has sufficient stability and reproducibility to be measured with behavioral-science methods, but the paper's own cited evidence repeatedly shows acute sensitivity to prompts, context, and model version. Section 2.1 reports \"LLM sensitivity to prompts\" [95, 171]; Section 6.2 notes that a single added emoji can significantly alter outputs [174]; Section 6.3 reports order effects in similarity judgments [156]; and Section 2.4 concedes that evaluations are limited in scale and scenario diversity. The paper never provides test-retest reliability data, variance decompositions, or a comparison of within-condition variability across paraphrases and model versions versus between-condition effects. Without such evidence, the \"emergent behavior\" the paradigm studies could be largely an artifact of prompt or version variance, undermining the brain-to-behavior analogy. This needs to be addressed explicitly, for example by adding a section on behavioral reliability, measurement invariance, and variance decomposition.","section":"Sections 2.1, 6.2, 6.3, 2.4"},{"comment":"The characterization of Mozikov et al. [107] is internally contradictory. Section 2.1 states that \"emotions can influence the strategic decision-making of LLMs in a manner similar to how they affect humans,\" while Section 2.2 states the same work shows \"many LLMs have emotional tendencies distinct from those of humans, making them potentially more rational,\" and Table 2 lists [107] under \"Decision making is not affected by emotions like humans.\" The same citation is used to support opposite conclusions in adjacent subsections, which undermines the reliability of the synthesis and must be corrected.","section":"Sections 2.1 vs 2.2; Table 2"},{"comment":"The paper transfers human behavioral theories—social cognitive theory in Section 2 and the Fogg Behavior Model in Section 5—to LLM-based agents without validating the transfer. The Fogg mapping (ability=pretraining, motivation=RL/fine-tuning, trigger=prompting) is presented as a post-hoc classification; the paper itself acknowledges in Section 7 that most methods \"were not originally developed with behavioral theory in mind\" and that the framework is retrospective. As presented, the framework cannot be falsified and does not generate novel predictions. The paper should clarify whether these are intended as testable theories or as organizing heuristics, and if the former, specify what empirical observations would disconfirm them.","section":"Sections 2 and 5"}],"minor_comments":[{"comment":"The first research direction heading contains a typo: \"ehavior\" should be \"behavior.\"","section":"Section 7"},{"comment":"The citation [79] for the \"dual-reward reinforcement learning architecture\" appears mismatched: the listed reference is \"Socially situated artificial intelligence enables learning from human interaction,\" which does not obviously describe the dual-reward RL method summarized in the text. Please verify and correct the citation.","section":"Section 5.2"},{"comment":"Reference [85] is missing a year and publication status; complete the bibliographic details.","section":"References"},{"comment":"Figure 4 has cramped and overlapping labels, making the Fogg-model mapping difficult to read; a cleaner layout would improve clarity.","section":"Figure 4"},{"comment":"The row describing \"ontological assumption\" under the behavioral perspective reads more like a normative stance than a descriptive contrast; consider rephrasing to maintain the table's analytic tone.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for an interdisciplinary behavioral AI venue, though the q-bio.NC subject category is unusual for this content. Several illustrative works cited in the synthesis—EconAgent [86], S3 [61], AgentSquare [133], and OpenCity [176]—share authors with the current paper; this is not improper, but a broader base of independent examples and a brief disclosure would strengthen the presentation. The main risk is that the \"necessary complement\" claim oversells the evidence; if the authors soften the claim, fix the internal contradictions, and address behavioral reliability explicitly, the paper could become a useful contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real contribution is the packaging: it takes the scattered 'agents behave like humans' literature and organizes it into a named paradigm with a three-level structure (individual, multi-agent, human-agent), a Fogg-style adaptation framework, and a mapping of responsible-AI principles onto behavioral properties. That package is genuinely useful for newcomers to the area, and the six future directions are concrete enough to be worth debating. The responsible-AI reframing—fairness, safety, interpretability, accountability, privacy as trajectories rather than static model properties—is the strongest and most thought-provoking part of the paper.\n\nThe soft spots are real but mostly fixable. The biggest mechanical problem is an internal contradiction in how the same reference is summarized: Table 2 and Section 2.1 say Mozikov et al. [107] shows emotions influence LLM decisions in a human-like way, while Section 2.2 says the same paper shows LLMs have emotional tendencies distinct from humans, making them more rational. Both cannot be accurate. For a paper whose entire job is synthesis, that is a sloppy error and it lowers confidence in other summaries.\n\nThe literature selection is also unsystematic, with no stated method. The heavy presence of the authors' own systems (EconAgent, S3, AgentSquare, OpenCity) as illustrative evidence is not disqualifying, but it should be disclosed. The central claim that the paradigm is a 'necessary complement' is asserted rather than demonstrated. That's acceptable in a position paper if the case is well made; here it is decent but not airtight.\n\nThe stress-test about prompt and version sensitivity lands, but I'd call it a serious challenge rather than a fatal one. The paper openly acknowledges that LLM outputs vary with prompts, order, and emojis—it even uses that sensitivity to justify a behavioral lens. What it doesn't do is address the reliability question head-on: can behavioral measurements be stable enough across paraphrases and model snapshots to ground a science? The proposed 'behavioral entropy' construct is a good start, but it's undeveloped, and the brain-to-behavior analogy is load-bearing without being defended.\n\nWho should read it: anyone entering the field of LLM-agent evaluation or looking for a taxonomy of agent behaviors. It deserves a serious referee; I'd want the contradiction fixed, the selection method disclosed, and a paragraph on measurement reliability before publication.","headline":"A useful, occasionally sloppy consolidation of the agent-behavior literature, whose central promise—measuring agents as behaviors rather than mechanisms—is real but not yet earned.","tokens_in":38601,"tokens_out":2080,"would_cite":true,"duration_ms":24530,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI agent behavior is not determined by the model alone but emerges from situated interaction, making it a scientific subject in its own right.","keywords":["AI agents","behavioral science","large language models","multi-agent systems","human-agent interaction","responsible AI","Fogg Behavior Model","machine behavior"],"falsifier":"A concrete observation that would settle the claim: if a fixed model and fixed task, with only incidental prompt phrasing or model version varied, produces behavior distributions that vary as much as, or more than, changes in the environmental and social variables the paradigm treats as formative, the situated-interaction foundation is undermined. Alternatively, demonstrating that agent behaviors observed in sandbox simulations do not predict behaviors of the same agents in deployment settings would falsify the paradigm's predictive promise.","tokens_in":37537,"feed_emoji":"🤖","tokens_out":3952,"duration_ms":39162,"temperature":0.7,"pith_summary":"This paper argues that the behavior of an AI agent is not fixed by the model inside it; it emerges from the agent's situated interaction with environments, other agents, and humans. It proposes AI Agent Behavioral Science as a new paradigm that studies what agents actually do, adapt, and become over time, treating the model as a substrate that enables but does not determine behavior. The paper systematizes findings across individual, multi-agent, and human-agent settings, and shows how responsible AI concerns such as fairness, safety, and privacy can be reframed as measurable behavioral properties. A sympathetic reader would care because, if the paradigm holds, evaluating and governing AI shifts from inspecting internal weights to observing and shaping behavioral trajectories.","feed_headline":"Behavior, not weights, predicts what AI agents do","feed_subtitle":"A new paradigm argues that fairness, safety, and trust should be measured as behavioral trajectories, not static model properties.","key_machinery":"The paradigm itself is the central object: AI Agent Behavioral Science, defined as 'the study of how AI agents act, adapt, and interact in situated contexts.' Carrying the argument are three organizing devices: the brain-to-action analogy, which licenses transferring behavioral-science methods to agents; the social cognitive theory triad of intrinsic attributes, environmental constraints, and behavioral feedback for individual behavior; and the Fogg Behavior Model (ability, motivation, trigger) for classifying adaptation techniques. These devices convert scattered empirical results into a structured scientific field with its own measurement, intervention, and theory-guided interpretation.","core_discovery":"The central claim is that AI agent behavior is a legitimate empirical subject in its own right. The paper states this as an ontological analogy: 'the model is to behavior what the brain is to action: a substrate that enables but does not determine.' From this, it follows that complex behaviors such as negotiation, deception, cooperation, and institutional formation are not properties of the LLM alone but products of the agentic system—memory, planning, tools, roles, feedback—embedded in context. The paper organizes the emerging literature into individual, multi-agent, and human-agent interaction layers, and interprets adaptation methods through the Fogg Behavior Model by mapping ability to pretraining, motivation to reward signals, and trigger to prompting. It then argues that responsible AI principles should be treated as dynamic, context-dependent behavioral attributes rather than static model properties, opening a research agenda on behavioral entropy, macro-level adaptation, artificial societies, and behavioral warning signs.","pith_inferences":["The logic of the paper implies that behavioral benchmarks should eventually replace or complement static model benchmarks for deployment decisions; a model card would be paired with a behavioral profile measured across contexts. This is our inference, not a claim the paper explicitly makes.","If behavioral science transfers to agents, then behavioral entropy and other summary statistics could function like temperament inventories for AI, enabling comparison across model versions and providers—an extension the paper proposes as a research direction, not an established result.","A testable consequence the authors leave implicit: two agents with the same underlying model but different memory, role, or feedback structures should diverge behaviorally over interaction rounds, and this divergence should be measurable and stable across repeated runs. A failure to observe such divergence would challenge the substrate-enables-behavior claim."],"forward_implications":["If the paradigm is right, evaluating an AI system means running behavioral experiments—observing trajectories over time and across contexts—rather than only auditing weights or static outputs.","Responsible AI metrics (fairness, safety, interpretability, accountability, privacy) would each need trajectory-level operationalizations, since one-shot assessments miss drift, deception, and feedback effects.","Adaptation research would be organized around behavioral levers—ability, motivation, trigger—so that prompting, fine-tuning, and reinforcement learning are seen as complementary ways to shape behavior, not competing paradigms.","Artificial societies made of agents become legitimate instruments for behavioral theory, allowing controlled, replicable, counterfactual experiments that are impossible with human subjects.","Hybrid human-agent teams and machine culture become objects of scientific study, with the field predicting that behavior, not architecture, determines collective outcomes."],"supporting_citations":[{"why":"Supplies the precedent of treating AI systems as empirical subjects of behavioral study, which the paper extends into a full paradigm.","marker":"[121]"},{"why":"Demonstrates behavioral-science tools applied to LLM preferences and traits, providing the key empirical basis for the transfer of methods.","marker":"[103]"},{"why":"Supplies the social cognitive theory taxonomy of intrinsic attributes, environmental constraints, and behavioral feedback used to organize individual agent behavior.","marker":"[13]"},{"why":"Supplies the Fogg Behavior Model whose ability–motivation–trigger structure organizes the paper's account of behavior adaptation.","marker":"[57]"},{"why":"Provides the sandbox evidence of emergent social behavior in generative agents, the motivating example for studying situated interaction.","marker":"[115]"},{"why":"Frames human-AI hybrid networks as complex systems with emergent dynamics, reinforcing the situated-context view.","marker":"[155]"},{"why":"Shows that AI systems participate in generating and transmitting cultural patterns, supporting the need for behavioral observation of agents in social ecosystems.","marker":"[17]"},{"why":"Defines the AI agent as an autonomous system that perceives and acts, the unit of analysis for the proposed paradigm.","marker":"[167]"}],"fun_headline_variants":["AI behavior is its own science, not just model internals","Context shapes AI agent behavior","Fairness and safety are behaviors, not model traits","Why AI agents need a behavioral science","Study AI behavior, not just weights"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paradigm assumes that human behavioral-science methods transfer to LLM-based agents, meaning that agent behavior has stable, context-dependent, causally structured regularities that hold beyond the specific sandbox where they were observed.","fun_headline_variants_meta":{"raw":{"variants":["AI behavior is its own science, not just model internals","Context shapes AI agent behavior","Fairness and safety are behaviors, not model traits","Why AI agents need a behavioral science","Study AI behavior, not just weights"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001274,"raw_usage":{"total_tokens":5200,"prompt_tokens":922,"completion_tokens":4278,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":4211}},"tokens_in":538,"tokens_out":4278,"duration_ms":35459,"temperature":1.0,"reasoning_tokens":4211,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:57:01.054850+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete observation that would settle the claim: if a fixed model and fixed task, with only incidental prompt phrasing or model version varied, produces behavior distributions that vary as much as, or more than, changes in the environmental and social variables the paradigm treats as formative, the situated-interaction foundation is undermined. Alternatively, demonstrating that agent behaviors observed in sandbox simulations do not predict behaviors of the same agents in deployment settings would falsify the paradigm's predictive promise.","supporting_citations":[],"review_version":1}