{"id":"f6908a96-c920-4b7c-abfe-b1645f277bf1","arxiv_id":"2507.05723","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review showing that current LLM use in agent-based modelling is concentrated in implementation, with untapped but risky potential across problem formulation, conceptualization, verification, validation, interpretation, and documentation.","lead":"This paper reviews how large language models are currently and could potentially be used across every stage of building agent-based simulations. It provides a structured map of opportunities, pitfalls, and mitigations for each phase of the modelling cycle.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Scopus-only corpus may skew the 91% implementation concentration; the paper's own Section 2 cites relevant non-coded work that could change the phase distribution.","rationale":"The paper's descriptive claim is clearly stated and internally consistent with its coding: 20 of 22 papers are coded as using LLMs in implementation. The methodology is transparent—full query in the annex, eligibility criteria, dual coding with follow-up discussion—and the forward-looking map is carefully hedged as potential uses. The weakness is external validity. A rapid review of a single database is a reasonable starting point, but the paper uses the resulting 22-paper corpus to conclude that the rest of the ABM cycle is 'not yet fully leveraged.' This is an inference from absence in a narrow corpus. The paper itself lists relevant related work in Section 2 (e.g., [26], [40], [38]) that did not enter the corpus, some of which discuss non-implementation phases; readers cannot tell whether these were excluded by the stated eligibility criteria or by database coverage. An expanded search plus recoding would settle whether the 91% concentration is robust. I do not see a reason to move beyond CONDITIONAL: the synthesis remains useful if framed as scoped to the reviewed corpus, but the empirical headline should not be generalized without this check.","tokens_in":12973,"tokens_out":8214,"duration_ms":89804,"concrete_test":"Re-run the rapid review across Web of Science, IEEE Xplore, and arXiv (including preprints mentioned in the paper's own Section 2), applying the same eligibility criteria, and recode the phase distribution. If the share of papers with LLM use concentrated in implementation remains at or above 80% and no non-implementation phase gains multiple papers, the concentration finding stands; if the share drops below roughly 70% or non-implementation phases become populated, the paper's 'unexploited phases' claim needs to be re-scoped or weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central descriptive claim—that current LLM use in ABM is concentrated in implementation (20/22, 91%) and that the rest of the modelling cycle remains unexploited—is only as strong as the 22-paper corpus from which it is drawn. The corpus comes from a Scopus-only rapid review with a March 2025 cutoff. Section 2 of the same paper cites relevant works that are not in the coded set, including arXiv preprints and non-Scopus-indexed venues (e.g., Siebers [40], Larooij and Törnberg [26], Polhill et al. [38]); several of these discuss LLM use in model design and other non-implementation phases. If such works, or similarly missed papers, had been included, the 91% concentration could be substantially lower. The inference from 'not found in Scopus' to 'phase unexploited in the field' is therefore fragile. Since the forward-looking map is motivated by this perceived gap, the skew propagates to the paper's main synthesis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a Rapid Literature Review (RLR) of 22 Scopus-indexed papers that describe implemented uses of large language models in agent-based modelling. It finds that 20 of the 22 papers (91%) use LLMs in the implementation phase, mostly as LLM-powered agents for reasoning, decision-making, or communication, with one paper focused on code generation and one on interpretation of results. The paper then develops a structured map of possible uses across the ABM cycle—problem formulation, system analysis, conceptualization, implementation, verification, validation, interpretation and communication, and documentation—pairing each phase with potential pitfalls and mitigations. The forward-looking map is based on monthly group discussions among the co-authors rather than on a systematic community survey, and the paper ends with a critical reflection on the symbolic versus data-driven tension between ABM and LLMs.","tokens_in":13168,"tokens_out":6987,"duration_ms":77652,"significance":"The paper's main value is its structured taxonomy of opportunities, challenges, and mitigations for LLM use across the entire ABM cycle, which is more comprehensive than prior surveys focused on single stages or single LLM capabilities. The descriptive RLR is transparently reported: the search query is given in the annex, inclusion/exclusion criteria are stated, dual coding with discussion is described, and the 22 coded papers are listed. The 91% implementation concentration, if reliable, is a clear and actionable observation for the community. The forward-looking sections are explicitly framed as critical reflection rather than empirical evidence, and the discussion of LLM-powered agents contains a substantive analysis of the tension between symbolic and data-driven AI. The principal limitations are the narrow empirical base for the concentration claim and the unquantified reliability of the phase coding.","major_comments":[{"comment":"The central descriptive claim that current LLM use in ABM is concentrated in implementation (20/22, 91%) depends entirely on the corpus retrieved from a single database (Scopus) with a March 2025 cutoff. Section 2 of the same paper cites relevant works that are not in the coded set, including pre-prints and works addressing model design, simulation tasks, and result interpretation (e.g., [14], [17], [40], [2], and [38]). If such works, or other papers not indexed in Scopus, had been included, the phase distribution could shift materially. The paper should either recompute the distribution after expanding the corpus to arXiv and other databases, or explicitly rescope the conclusion to \"the Scopus-indexed implemented-use corpus\" rather than claiming that \"the inherent potential of LLMs has not yet been fully leveraged throughout the ABM cycle\" (end of Section 5.1). This rescoping is load-bearing because the forward-looking map is motivated by the perceived gap.","section":"Section 4 / Annex / Section 5.1"},{"comment":"The reliability of the phase coding is asserted but not quantified. The method section says two coders \"checked each other's coding and discussed it extensively,\" but no codebook, disagreement rate, or intercoder agreement statistic is reported. Since the single most important number in the paper is the 20/22 phase classification, the absence of a reproducibility measure for that classification is a substantive gap. Please add a summary of coding disagreements and their resolution, or a quantitative reliability measure such as Cohen's kappa per phase code, and make the full codebook available in the annex if space permits.","section":"Section 4 / Table 1"},{"comment":"The reporting of the main count is internally ambiguous. The text says \"The most common use (n=20, 91%) involves implementation\" and then lists code generation [31] as the main focus of one paper and interpretation [30] as the main focus of another. However, Section 5.2 defines the Implementation phase as including code generation. The paper should clarify whether the 20 implementation papers include or exclude code-generation papers. If code generation is considered part of implementation, the count for that phase would be 21/22 rather than 20/22; if it is excluded, the definition of \"implementation\" in Section 5.1 differs from the one used in Section 5.2, which confuses the central statistic.","section":"Section 5.1 / Section 5.2 (Implementation)"}],"minor_comments":[{"comment":"The sentence listing limitations of previous studies contains an apparent contradiction: it says previous studies \"do not follow a straightforward structure\" and then says they \"follow a straightforward structure that is not specifically tailored for ABM.\" One of these clauses is likely a typo and should be corrected.","section":"Section 2"},{"comment":"The enumerated list of LLM-assisted verification tasks jumps from item (2) to item (5); the numbering should be corrected to (1), (2), (3), (4).","section":"Section 5.2 (Verification)"},{"comment":"The search query as printed contains unnatural spacing and line breaks (e.g., \"prompt e n g i n e e r i n g\"), which makes it difficult for readers to reproduce. A clean, copy-pasteable version of the query should be provided.","section":"Annex"},{"comment":"The paper reports that the RLR used a single database and a March 2025 cutoff but does not explicitly discuss the implications of these choices for the fast-moving LLM literature. A short limitation note, even one or two sentences, would help calibrate readers' expectations.","section":"Section 4"},{"comment":"The phrase \"and and\" appears in the acknowledgment for Vivek Nallur's grants; this should be corrected.","section":"Acknowledgments"},{"comment":"The forward-looking map is based on structured group discussions among the co-authors, but the paper does not describe the discussion protocol (e.g., number of sessions, how suggestions were aggregated, how disagreements were resolved). Adding a brief description would improve transparency for a contribution that is partly a collective expert opinion.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of a computational social science venue and the RLR is transparently reported. A note for the editor: the paper's organizational framework (the ABM cycle) draws on a co-author's prior work [41], and several prior studies cited in Section 2 involve co-authors of this paper. This is not circular, but the descriptive claim should be examined with the corpus-bias concern in mind. The issues raised in the major comments are fixable within the manuscript's scope, so I do not recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful synthesis—the first phase-by-phase map of where LLMs are and could be used across the ABM cycle—and the authors are transparent about their method. The main caveat is that the descriptive \"current use\" claim (91% implementation) rests on a Scopus-only rapid review of 22 papers, and the paper's own Section 2 lists relevant works that didn't make the coded corpus. That doesn't sink the paper, but it should be scoped more carefully.\n\nWhat's new: prior reviews focus on single phases or single LLM capabilities. This one walks through all phases (problem formulation, system analysis, conceptualization, implementation, verification, validation, interpretation, documentation), and for each gives concrete \"how\", \"pitfalls\", and \"mitigations\". That structure is usable for teaching, for tooling decisions, and for reporting standards like RAT-RS. The RLR is reported in enough detail to audit: query in annex, inclusion/exclusion criteria, dual coding with consistency checks, and the full list of 22 coded papers. That's more reproducible than most reviews in this space.\n\nSoft spots: the corpus is small and Scopus-only, with a March 2025 cutoff. The paper's own Section 2 cites works on model design and simulation tasks (e.g., [40], [14], [17]) that are not in the coded set. Some of those are excluded for good reasons (position papers, no implemented ABM), but at least a couple could plausibly have been included and would have shifted the phase distribution. So the \"91% implementation\" figure is best read as \"91% of the Scopus-indexed implemented papers we found,\" not \"91% of the field.\" The forward-looking map comes from group discussions among the co-authors; that's fine for generating hypotheses, but it shouldn't be mistaken for a community survey. The authors could have been more explicit about both limits in the conclusion, where they say \"the inherent potential of LLMs has not yet been fully leveraged.\" That's a reasonable interpretation, but it's a step beyond the evidence.\n\nOn balance, the paper deserves a serious referee. The structure and the candid reporting are solid; the main fixes are scoping the descriptive claim and maybe adding a sensitivity note about the search. I'd cite it as the standard reference for the ABM-cycle map.","headline":"A useful, transparent phase-by-phase map of LLM use in ABM, but the '91% implementation' concentration is a claim about a narrow Scopus-only corpus and should be scoped accordingly.","tokens_in":13730,"tokens_out":2918,"would_cite":true,"duration_ms":30754,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that LLM use in agent-based modelling has so far clustered almost entirely in the implementation phase, with 91% of reviewed papers using LLMs there, and that the rest of the modelling cycle remains largely unexploited.","keywords":["Large Language Models","Agent-Based Modelling","Modelling Cycle","Rapid Literature Review","Social Simulation","LLM-powered agents","problem formulation","verification and validation"],"falsifier":"A comparable systematic search across several other bibliographic databases using the paper's own inclusion criteria, finding more than a small number of implemented LLM uses in problem formulation, system analysis, or interpretation phases, would directly weaken the claim that 91% of current use is in implementation; the search is easy to run because the paper gives its full query in an annex.","tokens_in":12788,"feed_emoji":"🤖","tokens_out":5295,"duration_ms":52441,"temperature":0.7,"pith_summary":"The paper sets out to answer two questions: where and how are large language models actually used in agent-based modelling today, and where and how could they be used across the entire modelling cycle. A rapid review of 22 papers found that 91% of current uses sit in the implementation phase, mostly as LLM-powered agents that reason, deliberate, or communicate, while only two papers touch other phases. The remainder of the paper walks through the modelling cycle phase by phase, and for each phase lists concrete potential LLM uses, associated pitfalls, and mitigations. The authors conclude that LLMs are best seen not as suppliers of finished model components but as interpretive tools that provoke rethinking of system conceptualisations.","feed_headline":"LLMs in agent models: 91% of uses are in implementation","feed_subtitle":"A phase-by-phase map shows problem formulation, validation, and documentation are still open for LLM support.","key_machinery":"The organising device is the agent-based modelling cycle, taken from reference [41] and elaborated with reference [34]. The cycle divides modelling into problem formulation, system analysis, conceptualisation, implementation, verification, validation, interpretation and communication, and documentation. The paper uses this structure twice: first as a coding scheme for the literature review to locate where LLMs are used, and then as a template for a systematic opportunity map. The same structure also generates the pitfalls and mitigations, since each phase imposes different demands on text, transparency, and domain knowledge.","core_discovery":"The central discovery is empirical and structural. After coding 22 papers retrieved from a single literature database, the authors report that 20 of them (91%) use LLMs at the implementation stage, and that this use is almost always to power agents with reasoning, decision-making, or communication abilities. Only one coded paper focuses on code generation and one on interpreting model results. From this concentration the paper argues that the field has left most of the modelling cycle unexploited, and it offers a phase-by-phase map of opportunities, from problem formulation and system analysis through conceptualisation, verification, validation, interpretation, and documentation, each with risks and mitigations.","pith_inferences":["The empirical concentration may be partly a publication artefact: implementation with LLM agents is the most demonstrable and citable use, while LLM support for problem formulation or validation is harder to showcase in a conference paper.","The same single-database design, if applied to other simulation communities, might show a similar implementation-heavy pattern, and the paper's cycle framework could be reused for those cross-community audits.","A testable extension would be to benchmark an LLM-assisted modelling workflow against a traditional one across a full cycle, measuring time-to-model, number of errors, and interpretive quality; the paper does not report such measurements but its map implies the need for them.","The paper's mitigations imply new reporting infrastructure: a standard way to disclose LLM involvement per phase, which could be piloted in venues requiring structured documentation protocols."],"forward_implications":["If current use is this concentrated, then the largest set of untested LLM applications lies outside implementation, and the paper's phase-by-phase map functions as a research agenda.","Modelers who treat LLM outputs as provisional and triangulate with experts and data get a way to reduce the hallucination and bias risks the paper catalogues.","The paper's reframing, where LLMs act as interpretive provocateurs rather than model suppliers, changes what an LLM-assisted modelling workflow is for, shifting value toward problem framing and communication.","Documentation practices like the ODD protocol would need to record where and how LLMs assisted, a direct call the paper makes for implementation-phase code.","The duality of symbolic ABM logic and data-driven LLM semantics suggests the two paradigms can complement rather than replace each other."],"supporting_citations":[{"why":"Supplies the modelling-cycle structure that organises the literature-review coding scheme and the opportunity map.","marker":"[41]"},{"why":"Provides a complementary method for developing agent-based models, used to shape the phase descriptions.","marker":"[34]"},{"why":"Defines the rapid literature review phases followed to identify, select, and codify the 22 papers.","marker":"[7]"},{"why":"Is the ODD protocol referenced when discussing conceptualisation and documentation support.","marker":"[19]"},{"why":"Is the RAT-RS reporting standard that the paper suggests extending with a column documenting LLM use.","marker":"[1]"}],"fun_headline_variants":["LLM use in agent-based models: 91% stuck in implementation","Agent-based modelling leans on LLMs, but only for implementation","Review maps LLM roles across modelling cycle, spots gaps","In agent models, LLMs power agents but skip validation and docs","Most LLM use in agent modelling happens at implementation only"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole analysis rests on a literature review that searched one database with one query as of March 2025, together with the authors' own group discussions for the forward-looking claims; if the search missed substantial work, or the group's judgment skews toward their own interests, both the 91% concentration result and the priority map could shift.","fun_headline_variants_meta":{"raw":{"variants":["LLM use in agent-based models: 91% stuck in implementation","Agent-based modelling leans on LLMs, but only for implementation","Review maps LLM roles across modelling cycle, spots gaps","In agent models, LLMs power agents but skip validation and docs","Most LLM use in agent modelling happens at implementation only"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000717,"raw_usage":{"total_tokens":3144,"prompt_tokens":791,"completion_tokens":2353,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":407,"completion_tokens_details":{"reasoning_tokens":2266}},"tokens_in":407,"tokens_out":2353,"duration_ms":15708,"temperature":1.0,"reasoning_tokens":2266,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:18:55.568127+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A comparable systematic search across several other bibliographic databases using the paper's own inclusion criteria, finding more than a small number of implemented LLM uses in problem formulation, system analysis, or interpretation phases, would directly weaken the claim that 91% of current use is in implementation; the search is easy to run because the paper gives its full query in an annex.","supporting_citations":[{"cited_title":"Simulating social complexity: a handbook pp","cited_arxiv_id":null,"evidence_quote":"Supplies the modelling-cycle structure that organises the literature-review coding scheme and the opportunity map."},{"cited_title":"In: 2011 international conference on networking, sensing and control","cited_arxiv_id":null,"evidence_quote":"Provides a complementary method for developing agent-based models, used to shape the phase descriptions."},{"cited_title":"https://doi.org/10.1007/978-1-0716-1566-9","cited_arxiv_id":null,"evidence_quote":"Defines the rapid literature review phases followed to identify, select, and codify the 22 papers."},{"cited_title":"Ecological modelling221(23), 2760– 2768 (2010)","cited_arxiv_id":null,"evidence_quote":"Is the ODD protocol referenced when discussing conceptualisation and documentation support."},{"cited_title":"In- ternational Journal of Social Research Methodology25(4), 517–540 (2022)","cited_arxiv_id":null,"evidence_quote":"Is the RAT-RS reporting standard that the paper suggests extending with a column documenting LLM use."}],"review_version":1}