{"id":"3b3854f1-cf81-4a26-9e99-80fcd40b8083","arxiv_id":"2412.07788","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper reviews human behavior simulation research by behavior type, objective, and methodology, and identifies open problems such as multi-behavior joint simulation and LLM-driven agents.","lead":"This paper is a survey that organizes research on human behavior simulation into cognitive, physiological, social, and economic categories, and compares knowledge-driven, data-driven, and LLM-based methods. A generalist might read it to get a structured map of a field that supports recommendation systems, autonomous driving, and social media governance.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's taxonomy and selection rule are applied inconsistently, so the gap analysis that forms the paper's main contribution is not reproducible from the stated method.","rationale":"The reader's verdict is ACCEPT with moderate confidence, and the reader identifies the unstated selection rule as the weakest assumption. I agree that coverage selection is the critical vulnerability and that the paper is broadly useful as a reference. My stress test sharpens the reader's concern into a concrete, textually located correctness risk: the taxonomy and the exclusion rule are applied inconsistently within the paper itself, not merely left unspecified. The Entertainment row, the Section 3.5 garbled summary, and the unsupported NLP-maturity sentence are internal inconsistencies in the exact sections that establish the survey's main deliverable. Because the gap analysis is the paper's stated contribution, these inconsistencies make the central claim of systematic, comprehensive mapping not fully reproducible. This is a conditional-acceptance issue rather than a rejection: the survey's organization is plausible and its coverage is broad, but the authors should state the inclusion/exclusion protocol, fix the Entertainment labeling, and provide the coding of references into the taxonomy. I partially agree with the reader because they see the selection issue as primarily editorial, while I see it as load-bearing for the gap analysis; however, I do not dispute the overall value of the survey as a reference contribution.","tokens_in":33142,"tokens_out":2856,"duration_ms":25243,"concrete_test":"Recode the full reference list into the four-category, two-objective, micro/macro taxonomy using only the Section 2 definitions and the stated exclusion rule as the codebook, without looking at Tables 1 and 2. Then compute per-cell agreement with Tables 1 and 2, and identify which cells change from blank to filled or filled to blank. If the disagreement is concentrated in cells that drive the Section 5 open-problem claims (e.g., microscopic social behavior, entertainment, cognitive decision-making), then the gap analysis is not robust to the paper's own taxonomy and the 'comprehensive' claim should be softened to 'selective review.'","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is that it provides a comprehensive, systematic map of human behavior simulation whose main output is a gap analysis (Tables 1 and 2) feeding the open-problem discussion in Section 5. This claim depends on the taxonomy being applied consistently and on the stated exclusion rule being well-defined. Both are load-bearing and both are unstable. In Section 3 the authors write 'we do not intend to be exhaustive... those behaviors with few works simulating it are not included,' but no operational rule is given for what counts as a behavior, when a behavior has 'few works,' or how a work is assigned to a cell. The instability is visible in the text itself. Section 3.4.2 says 'Since there are few works simulating offline entertainment behaviors, we don't include them in this survey,' yet Table 2 lists an Entertainment row with six citations, without any note that the row covers only online entertainment. Similarly, Section 3.5 contains the garbled sentence 'simulation studies of social, economic, and social behavior are more developed' and an orphaned definitional sentence about physiological behaviors, exactly in the passage that is supposed to summarize the taxonomy-driven findings. The Section 3.5 text directly undercuts the Section 2 claim that cognitive behavior is the foundation of the other three: Section 3.5 states 'With the maturity of natural language processing technology, some researchers suggest that it is time to improve the simulation accuracy for better decision-making [100, 101],' but refs [100, 101] are hate-speech-propagation and opinion-dynamics works, not LLM-based work, so the sentence does not support the surrounding argument about NLP maturity. These are not cosmetic typos: the central deliverable is the blank-cell pattern in Tables 1 and 2, and that pattern shifts if the inclusion rule is applied differently.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of human behavior simulation. It proposes a taxonomy of four behavior types (cognitive, physiological, social, economic), two simulation objectives (scientific discovery and decision-making), two perspectives (microscopic and macroscopic), and three methodological families (knowledge-driven, data-driven, and knowledge-data co-driven). It catalogs representative works in two tables, discusses canonical methods (social force model, percolation, opinion dynamics, RNN, GCN, GAN, RL, GAIL, knowledge-infused learning, and LLM agents), and concludes with open problems. The central claim is that this is the first updated, comprehensive cross-disciplinary survey of human behavior simulation and that its tables reveal the gaps and opportunities for high-impact research.","tokens_in":33362,"tokens_out":5245,"duration_ms":46582,"significance":"The manuscript is useful as a structured map of a fragmented field. Its strengths are the breadth of the reference set, the explicit four-behavior taxonomy, the two-objective organization of Tables 1 and 2, and the discussion of LLM-based agents as a new methodological wave. It does not present derivations or new empirical results, and its contribution is organizational rather than falsifiable. If the taxonomy and table-construction rules are made precise, the paper could serve as a common vocabulary for researchers in transportation, computational social science, and recommender systems.","major_comments":[{"comment":"The inclusion rule for the survey is not operational. The text says 'we do not intend to be exhaustive... those behaviors with few works simulating it are not included,' but it never defines what counts as a behavior, how many works qualify as 'few,' or how a given work is assigned to a behavior row, objective column, and perspective cell. Because the blank cells in Tables 1 and 2 are the evidence for the gap analysis in Section 5, this selection rule is load-bearing. Please provide the operational criteria used to build the tables, or explicitly reframe the tables as an illustrative sample rather than a comprehensive gap map.","section":"Section 3 and Tables 1-2"},{"comment":"The summary sentence 'simulation studies of social, economic, and social behavior are more developed' is internally inconsistent (social is listed twice), and the intended contrast with the preceding cognitive-behavior sentence is unclear; presumably 'physiological' was meant. This sentence is the section's main takeaway and should be corrected. The following orphaned sentence defining physiological behaviors is also misplaced here; it belongs in Section 2, where the taxonomy is introduced.","section":"Section 3.5"}],"minor_comments":[{"comment":"In the Work row, reference [160] appears twice; the duplicate should be removed.","section":"Table 2"},{"comment":"There are typos in this section: 'utility fuctions' should be 'utility functions' and 'netowks' should be 'networks'.","section":"Section 3.2.1"},{"comment":"The phrase 'By combining these two processes, In this way, GAIL can learn...' is grammatically tangled and should be rewritten.","section":"Section 4.2.5"},{"comment":"The sentence beginning 'A pedestiran wants to reach...' contains a typo; also, the final sentence of the paragraph has awkward punctuation after 'calibrate'.","section":"Section 4.1.1"},{"comment":"The caption says that darker shading means more related works, but no quantitative or ordinal scale is given; please ensure the shading is legible in grayscale print or replace it with an explicit count or rank.","section":"Tables 1 and 2"},{"comment":"The abstract and Section 1 describe the survey as 'comprehensive,' while Section 3 states that the authors do not intend to be exhaustive; the wording should be reconciled so readers know the intended scope.","section":"Abstract and Section 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a CS.HC venue as a survey, and the structural issues are fixable; I do not see grounds for rejection. The main point to check in revision is whether the tables can be reconstructed from an explicit protocol; if not, the authors should soften the 'comprehensive' language. The self-citations are appropriate for a survey covering the authors' own prior crowd-simulation and LLM-agent work and do not raise a circularity concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — this survey is worth sending to referees, but I'd brace them for a sloppy middle section. The genuinely new thing is the organization by simulation target — cognitive, physiological, social, economic behavior — crossed with objective and micro/macro perspective. That framing is useful for people entering the area, and the LLM-era coverage is current. The method discussion (knowledge-driven, data-driven, co-driven, LLM agents) is competent and gives a fair map of the toolbox. I agree with the reader's ACCEPT verdict; this is a legitimate review contribution.\n\nThe soft spots are mostly editorial, but they cluster in the section that is supposed to carry the paper's analytic weight. Section 3.5 contains a garbled sentence ('social, economic, and social behavior') and an orphaned definition of physiological behavior. The exclusion rule is stated but underspecified — 'behaviors with few works simulating it are not included' without a count or protocol — so the blank cells in Tables 1 and 2 are a heuristic pattern, not a reproducible gap analysis. The paper should say that plainly. There is also a genuine mis-citation: the sentence about NLP maturity cites [100,101], which are hate-speech and opinion-dynamics papers, not LLM/NLP-technique papers. That should be fixed.\n\nWhere I part from the stress-test note: I don't think the instability of the selection rule is load-bearing enough to reject. The tables are explicitly non-exhaustive, and the qualitative conclusions survive even if a row boundary shifts. The Entertainment row is ambiguous because the text says offline entertainment is excluded while the table just says 'Entertainment,' but the cited works are uniformly online, so that inconsistency is cosmetic. The self-citation rate is noticeable but not abusive; the authors cite their own relevant crowd-simulation and LLM-agent work where appropriate.\n\nFor peer review: yes, send it. A careful referee should ask for a tightened Section 3.5, a one-paragraph operational definition of inclusion, and a fix to the [100,101] citation. The intended audience — grad students and practitioners looking for a map of the field — will get real value even with those flaws.","headline":"A useful but uneven survey: the target-based taxonomy is the real contribution, and the gap tables should be read as a heuristic, not a reproducible result.","tokens_in":33956,"tokens_out":2434,"would_cite":true,"duration_ms":23808,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey organizes the young field of human behavior simulation into four behavior types and two objectives, then maps which method combinations are mature and which are open.","keywords":["human behavior simulation","survey","taxonomy","large language models","agent-based modeling","social force model","opinion dynamics","decision-making"],"falsifier":"Run a systematic literature search for each cell in Tables 1 and 2, for example microscopic emotion simulation for decision-making or macroscopic creativity simulation. If substantial existing work appears in cells the survey marks empty, the comprehensiveness claim fails; if the blank cells stay blank, the gap analysis holds.","tokens_in":45,"feed_emoji":"🧠","tokens_out":4234,"duration_ms":105969,"temperature":0.7,"pith_summary":"This survey argues that the scattered field of human behavior simulation can be usefully organized by what is being simulated rather than by method or discipline. It sorts research into four behavior families — cognitive, physiological, social, and economic — and records, for each family, whether simulation serves scientific discovery or decision-making, at microscopic or macroscopic scale. The result is a maturity map: physiological behavior is the most studied and most method-rich, cognitive behavior is the least, and large language model agents are emerging as a new modeling paradigm across all four families. The paper's value is the map itself: it gives researchers a shared vocabulary and shows which combinations of behavior, objective, and method are crowded and which are empty. If the map is accurate, it points newcomers toward open problems such as joint multi-behavior simulation, multi-scale simulation, and simulation under abnormal conditions.","feed_headline":"Four behavior types map every human simulation approach","feed_subtitle":"The survey shows physiological simulation is mature, cognitive simulation is open, and LLM agents are the new frontier.","key_machinery":"The organizing device is a two-dimensional taxonomy: four behavior types (cognitive, physiological, social, economic) crossed with two objectives (scientific discovery vs. decision-making) and two perspectives (microscopic individual vs. macroscopic group). The paper presents this as Tables 1 and 2, with shading showing the number of works per cell. The second machinery is a methodological trichotomy — knowledge-driven, data-driven, and knowledge-and-data co-driven models — which the paper uses to explain how each behavior family has evolved and to locate LLM-based agents as the newest co-driven paradigm. This taxonomy is what carries the argument: it is the framework that makes the gap analysis possible.","core_discovery":"The paper's central claim is that it provides the first systematic survey to classify human behavior simulation by simulation target. It defines a taxonomy of four behavior types — cognitive behavior (reasoning, emotion, creativity, politico-religious behavior), physiological behavior (movement, driving, transitions, resource usage), social behavior (connection formation, influence, cooperation and competition), and economic behavior (work, entertainment, market) — and a two-part objective split: simulation as a scientific tool for understanding behavior, and simulation as an environment for decision-making. Across these categories it identifies a consistent methodological progression: knowledge-driven models (social force, percolation, opinion dynamics) come first, data-driven models (RNN, GCN, GAN, RL, GAIL) follow, and knowledge-and-data co-driven models, including LLM-based agents, are the current frontier. The survey claims this perspective reveals that physiological behavior simulation is the most mature, cognitive behavior simulation is the least data-driven and most idealized, and the most valuable future problems are joint simulation of multiple behaviors, multi-scale simulation linking individual and aggregate levels, and simulation of abnormal scenarios.","pith_inferences":["The taxonomy could be tested bibliometrically: if the survey's selection is representative, a full-text search of the four behavior terms should find the same maturity ranking, with physiological ahead of social and economic, and cognitive last.","The paper's frontier claims imply that LLM-based simulation, while flexible, carries unresolved efficiency and robustness costs; a natural extension is evaluating LLM agents against existing car-following or opinion-dynamics benchmarks to measure their accuracy gain over knowledge-driven baselines.","The blank-cell argument suggests that new research effort in cognitive behavior simulation would face a relatively open field, but the paper does not quantify publication volume, so that openness is an inference from the shading tables rather than a counted claim."],"forward_implications":["Researchers can transplant mature methods from physiological behavior (e.g., social force models, trajectory generation) to less mature cognitive and economic domains, since the survey identifies the analogous problem formulations.","Decision-making applications can adopt LLM-based agents where data is scarce, since the survey shows LLM heterogeneity and reasoning working across all four behavior families.","The blank cells in Tables 1 and 2 become a concrete research agenda, for example microscopic creativity simulation for decision-making, which the survey marks as open.","The consistency of methodological progression across disciplines suggests that a new behavior family will first receive knowledge-driven models, then data-driven models, then co-driven models."],"supporting_citations":[{"why":"Supplies the canonical knowledge-driven social force model used throughout physiological movement simulation and later co-driven work.","marker":"[53]"},{"why":"Defines the percolation and influence-spread framing that cognitive and social diffusion simulations build on.","marker":"[214]"},{"why":"Provides the foundational DeGroot opinion dynamics model that the cognitive and politico-religious simulation line extends.","marker":"[221]"},{"why":"Supplies agent-based modeling as the core tool for social and economic simulation in the survey's method taxonomy.","marker":"[29]"},{"why":"Surveys LLM-empowered agent-based modeling, the paradigm the paper identifies as the newest co-driven direction.","marker":"[20]"},{"why":"Provides the generative agents system that demonstrates LLM-based human behavior simulation in a virtual town.","marker":"[23]"},{"why":"Prior crowd simulation review used to establish physiological behavior as the most mature simulation domain.","marker":"[16]"}],"fun_headline_variants":["Four behavior types map every human simulation approach","Survey classifies human behavior simulation into four types","Physiology mature, cognition open: survey maps simulation landscape","LLM agents at frontier of human behavior simulation survey","From physics models to LLM agents: behavior simulation survey"],"cache_read_input_tokens":36096,"weakest_assumption_plain":"The survey's whole map rests on the selection of cited papers being representative; the authors say they are not exhaustive and exclude behaviors with few simulated works, so if the selection skews toward certain disciplines, the gap analysis would mislead.","fun_headline_variants_meta":{"raw":{"variants":["Four behavior types map every human simulation approach","Survey classifies human behavior simulation into four types","Physiology mature, cognition open: survey maps simulation landscape","LLM agents at frontier of human behavior simulation survey","From physics models to LLM agents: behavior simulation survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1296,"prompt_tokens":942,"completion_tokens":354,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":279}},"tokens_in":558,"tokens_out":354,"duration_ms":3967,"temperature":1.0,"reasoning_tokens":279,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:28:54.283374+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a systematic literature search for each cell in Tables 1 and 2, for example microscopic emotion simulation for decision-making or macroscopic creativity simulation. If substantial existing work appears in cells the survey marks empty, the comprehensiveness claim fails; if the blank cells stay blank, the gap analysis holds.","supporting_citations":[{"cited_title":"Maximizing the spread of influence through a social network","cited_arxiv_id":null,"evidence_quote":"Defines the percolation and influence-spread framing that cognitive and social diffusion simulations build on."}],"review_version":1}