REVIEW 4 major objections 5 minor 103 references
Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A generative agent can pass behavior-equivalence tests yet follow artificial decision processes.
desk verdict A real dual-validation framework and a genuinely new LLM-vs-human comparison, but the process-fidelity score is a sign-blind significance counter, so the headline paradox needs a robustness check before it carries the weight the paper puts on it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The dual-validation framework is the paper's central mechanism. It pairs Two One-Sided Tests (TOST) for surface-level equivalence — testing whether each LLM's mean response on tip change, joint satisfaction, and differential satisfaction lies within a pre-specified margin of the human mean — with multi-group Structural Equation Modeling (SEM) for decision-process validation. The SEM specifies a moderated mediation model: service outcome affects tip change, tip adjustability moderates that path, tip visibility moderates the path from tip change to dyadic satisfaction, and both satisfaction outcomes are direct products of service outcome and tip change. Process fidelity is operationalized as t
What would settle it
Run the same dual validation in a different dyadic LSCM setting, such as buyer-supplier negotiation, where the human SEM is independently validated with think-aloud process data. If a model passes TOST equivalence on all outcome measures and also matches all ten human pathways, then the equivalence-versus-process paradox is not a general property of LLM agents.
Extended reading notes
Core claim
The paper's central claim is that surface-level behavioral equivalence does not guarantee that LLMs replicate human decision-making processes in GABM applications. This is demonstrated as an equivalence-versus-process paradox: some LLMs produce outcomes statistically indistinguishable from humans while employing artificial internal decision pathways, and models with lower outcome fidelity can have more human-like processes. The evidence comes from a 4×2×2 factorial experiment in food delivery, where six LLMs (480 dyads each) and 477 human dyads responded to identical vignettes. Surface equivalence was assessed with TOST tests at a ±0.2 SD margin; process fidelity was scored by how many of ei
Load-bearing premise
The entire process-fidelity comparison rests on the human structural equation model being the correct and complete account of decision making in this dyad; if that model is misspecified, the pathway-matching scores and the paradox they reveal are built on sand.
Editorial extensions
If this is right
- Researchers running GABMs must declare which validation level their question demands; passing one level tells you nothing about the other.
- LLM selection for operational simulations should be guided by the task: surface equivalence is enough for outcome prediction, while mechanism-oriented theory and policy analysis require process validation.
- The 2025 model generations did not beat the 2024 generations on human behavioral equivalence in this context, so 'state of the art' status is not a proxy for behavioral fidelity.
- If LLMs hold artificial decision pathways despite correct outputs, then simulated emergent phenomena — even when statistically plausible — may be attributed to the wrong causal mechanisms.
- The study provides a template for validation in other dyadic LSCM contexts, with method menus for surface and process validation in each domain.
Reading between the lines
- The equivalence-versus-process paradox may be a general property of LLM agents, not a quirk of food delivery: because model training optimizes text likelihood rather than psychological mechanism, outcome matching and process matching are decoupled incentives. Testing this in buyer-supplier or shipper-carrier dyads would show whether the paradox generalizes.
- The process-fidelity score treats all ten pathways as equally important, so rankings could shift with a weighted scoring that emphasizes theoretically central paths (service outcome to satisfaction) over secondary ones; this is my own methodological suggestion, not the paper's.
- The four models with 8/10 pathway fidelity each still deviated on two paths, so even the 'best' process match is partial; exact coefficient equality, not just significance matching, is a stricter bar the paper does not apply.
- The finding that newer models were less surface-equivalent could reflect changes in instruction-following or response calibration rather than true behavioral drift; a longitudinal study with matched prompts would separate these.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a dual-validation framework for generative agent-based models (GABMs) in logistics and supply chain management: surface-level behavioral equivalence testing (TOST) and decision-process validation (SEM). Using a 4×2×2 vignette experiment on food-delivery customer–worker dyads, the author compares six LLMs against 957 human participants (477 dyads) and reports that GPT-4o passes surface-level equivalence on all three behavioral measures but has the lowest process fidelity (5/10 SEM pathway matches), while GPT-4.1, Sonnet 3.5, Sonnet 4, and Mistral Medium 3 match 8/10 human pathways despite weaker surface equivalence. The paper interprets this as an 'equivalence-versus-process paradox' and argues that both validation levels are necessary in GABM development.
Significance. The paper addresses a timely and important gap: systematic validation of LLM-based agents in behavioral LSCM research. The human baseline is large (477 dyads) and the experimental design is careful, including parallel vignettes and response formats for humans and LLMs. The proposed dual-validation idea is sensible and method-agnostic, and the study provides one of the first empirical demonstrations that surface-level output equivalence can diverge from decision-process fidelity. However, the central paradox rests on a process-fidelity metric that is a sign-blind, threshold-dependent count of significant pathways. Because the metric treats opposite-sign coefficients as matches and ignores effect sizes, the headline ranking and the claimed paradox are not as robust as the manuscript suggests. The human indirect effects are all non-significant, so part of the 'fidelity' score amounts to matching null patterns. These issues are fixable with additional analyses, but they are load-bearing for the main contribution.
major comments (4)
- [§4.5.4 and Table 5; Appendix C] The process-fidelity score counts a pathway as 'matched' if the LLM's p-value is on the same side of 0.05 as the human p-value, ignoring sign and magnitude. This is not a minor coding choice: the central paradox is operationalized through this metric. For example, the human Joint satisfaction ~ Tip change path is β=+0.051 (p<0.001, Table C.1), while GPT-4.1's path is β=-0.040 (p<0.001, Table C.3). Both are counted as significant matches, contributing to GPT-4.1's 8/10 fidelity score, even though the LLM reproduces the human path with the opposite sign. Similarly, Differential satisfaction ~ Service outcome has human β=0.936 (Table C.1) vs GPT-4.1 β=0.350 (Table C.3), yet both count as matches. A model that replicates human decisions through an opposite psychological mechanism can therefore be labeled 'high process fidelity.' I recommend re-scoring with sign agreement required and/or a st
- [§4.5.4, Table C.8] The human model shows non-significant indirect effects on both joint and differential satisfaction (p = 0.606 and p = 0.287). The process-fidelity metric therefore rewards LLMs for reproducing null indirect effects. But matching a null pattern is much weaker evidence of authentic process replication than matching a well-powered positive effect. The human sample (477 dyads) is not obviously underpowered for the direct effects, but the indirect effects are small and the bootstrap CIs are wide; the test may simply be insensitive. The paper should report the bootstrap power or equivalence bounds for these indirect effects, or at minimum explicitly acknowledge that the two indirect-effect 'matches' are matches to a null. As it stands, the process-fidelity ranking partly depends on LLMs correctly not showing something that the human data cannot strongly demonstrate.
- [§4.5.1 and Table 4] The TOST equivalence margin of ±0.2 SD is introduced without justification or sensitivity analysis. This margin is not merely a statistical detail: the surface-level ranking is the first pillar of the claimed paradox. For example, GPT-4o is the only model with equivalence on all three measures, but the p-values for tip change and differential satisfaction are close to 0.02 and 0.004, respectively. If a tighter margin (e.g., ±0.1 SD) were used, some of these equivalences might disappear; if a looser margin were used, more models would pass and the surface-level hierarchy would change. I request a rationale for ±0.2 SD (e.g., prior literature or practical relevance) and a sensitivity analysis over a plausible range of margins, reporting the resulting surface-level rankings and the implications for the paradox.
- [§4.5.4, multi-group SEM] The process-level validation compares each LLM's p-value pattern to the human p-value pattern, but it does not formally test whether the LLM coefficients differ from the human coefficients. A more rigorous multi-group approach would estimate a constrained model (e.g., fixing paths to the human estimates) and test the chi-square difference or use an equivalence test on individual coefficients. As written, the method treats a non-significant p-value in an LLM as a 'match' even when the coefficient estimate is far from the human estimate (e.g., Differential satisfaction ~ Service outcome: human 0.936 vs Mistral Medium 3 0.20, both significant, counted as a match). The paper should either implement formal tests of coefficient equality or clearly frame the process-fidelity score as a descriptive significance-pattern metric, not as evidence of process equivalence.
minor comments (5)
- [Table 4] The p-values in Table 4 are formatted inconsistently (e.g., 0.02, 0.00, 1.00, 0.34) and some are rounded to two decimals in a way that obscures exact values. Please use a consistent three-decimal format and mark values below 0.001 as '<0.001'.
- [§4.4.6] The 'white text haiku' quality-control measure deserves a clearer description. As written, it is not obvious how requesting a haiku in white text detects AI-generated responses. Clarify the mechanism or remove this sentence.
- [Table 3] Means and standard deviations in Table 3 are typeset with unusual spacing (e.g., '3 .61' instead of '3.61'). This appears to be a rendering issue, but it makes the table harder to read. Please fix the formatting.
- [§4.5.2] The sentence 'Three dyads were removed due to three human participants failing quality checks, making that participant's randomly assigned dyad unusable' is slightly confusing: it implies a 1:1 mapping from participant failure to dyad removal. Clarify how dyads were defined and how a single failed participant invalidated a dyad.
- [§5.3] The limitation about static vignettes is useful, but it could be expanded to note that the LLM agents do not interact with each other; they respond to a fixed scenario. This is relevant because the paper describes 'dyadic interactions' and 'emergent' behavior, but the current implementation is a sequential vignette response, not a multi-turn interaction. Please clarify the scope.
Circularity Check
No significant circularity: the validation compares independent LLM outputs against an external human SEM baseline; the sign-blind pathway metric is a validity concern, not a circular derivation.
full rationale
The paper's central claim is an empirical comparison, not a derivation. The human SEM (Table C.1) is estimated from 957 human participants; each LLM SEM (Tables C.2–C.7) is estimated from 480 dyads per model. The process-fidelity score counts significance matches across 8 direct paths and 2 indirect effects (Table 5). No parameter is fitted from LLM data to predict a human quantity; TOST equivalence margins (±0.2SD) are defined from human means, which is standard practice. The paper's self-citations (Castillo et al. 2018, 2021, 2022; Saunders et al. 2025) are domain literature and are not load-bearing for the validation framework or for the equivalence-versus-process claim. The sign-blindness of the pathway scoring (e.g., human Joint satisfaction ~ Tip change β=+0.051, p<0.001 vs GPT-4.1 β=−0.040, p<0.001, both counted as matches) is a substantive construct-validity limitation of the process-fidelity metric, but it does not make the comparison circular: the human baseline is independent, and the contradiction between surface equivalence and process fidelity is an observed empirical pattern, not a definitional identity. The paper's own limitation statements (static vignettes, aligned dyadic experiences) concern external validity, not circularity. Therefore no circular step is present.
Assumptions & free parameters
free parameters (4)
- TOST equivalence margin (±0.2 SD) =
0.2 SD
- SEM pathway set (8 direct paths + 2 indirect effects) =
10 pathways
- Bootstrap replications (5,000) =
5000
- LLM temperature and top_p (0.7, 0.95) =
temperature=0.7, top_p=0.95
assumptions (5)
- domain assumption The human SEM estimated from 477 dyads is the correct baseline representation of dyadic decision making in this context.
- domain assumption Surface-level equivalence is appropriately tested at the mean level with TOST over three outcome variables.
- domain assumption LLM responses at temperature 0.7 with identical prompts are valid individual-level behavioral samples.
- domain assumption The vignette-based paradigm captures the decision process of real dyadic food delivery encounters.
- domain assumption Structural zeros imputed for tip change in non-adjustable conditions do not distort the SEM estimates.
Cite this review
Pith. "Pith review of Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research." pith.science (2026). https://pith.science/paper/IICMEIL4
@misc{pith2026250820234,
author = {Pith},
title = {Pith review of: Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research},
year = {2026},
howpublished = {\url{https://pith.science/paper/IICMEIL4}},
note = {Machine review of arXiv:2508.20234}
}
read the original abstract
Generative Agent-Based Models (GABMs) powered by large language models (LLMs) offer promising potential for empirical logistics and supply chain management (LSCM) research by enabling realistic simulation of complex human behaviors. Unlike traditional agent-based models, GABMs generate human-like responses through natural language reasoning, which creates potential for new perspectives on emergent LSCM phenomena. However, the validity of LLMs as proxies for human behavior in LSCM simulations is unknown. This study evaluates LLM equivalence of human behavior through a controlled experiment examining dyadic customer-worker engagements in food delivery scenarios. I test six state-of-the-art LLMs against 957 human participants (477 dyads) using a moderated mediation design. This study reveals a need to validate GABMs on two levels: (1) human equivalence testing, and (2) decision process validation. Results reveal GABMs can effectively simulate human behaviors in LSCM; however, an equivalence-versus-process paradox emerges. While a series of Two One-Sided Tests (TOST) for equivalence reveals some LLMs demonstrate surface-level equivalence to humans, structural equation modeling (SEM) reveals artificial decision processes not present in human participants for some LLMs. These findings show GABMs as a potentially viable methodological instrument in LSCM with proper validation checks. The dual-validation framework also provides LSCM researchers with a guide to rigorous GABM development. For practitioners, this study offers evidence-based assessment for LLM selection for operational tasks.
Figures
Reference graph
Works this paper leans on
-
[1]
K., Lemarinier, P., and O’Hare, G
Abar, S., Theodoropoulos, G. K., Lemarinier, P., and O’Hare, G. M. (2017). Agent Based Modelling and Simulation tools: A review of the state-of-art software. Computer Science Review , 24:13--33
2017
-
[2]
and Meloy, M
Abbey, J. and Meloy, M. (2017). Attention by design: Using attention checks to detect inattentive respondents and improve data quality. Journal of Operations Management , 53-56
2017
-
[3]
C., and Sinchaisri, W
Allon, G., Cohen, M. C., and Sinchaisri, W. P. (2023). The Impact of Behavioral and Economic Drivers on Gig Economy Workers . Manufacturing & Service Operations Management , 25(4):1376--1393. Publisher: INFORMS
2023
-
[4]
and Pal, R
Altay, N. and Pal, R. (2014). Information Diffusion among Agents : Implications for Humanitarian Operations . Production and Operations Management , 23(6):1015--1027
2014
-
[5]
Ambra, T., Caris, A., and Macharis, C. (2021). Do You See What I See ? A Simulation Analysis of Order Bundling within a Transparent User Network in Geographic Space . Journal of Business Logistics , 42(1):167--190. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jbl.12237
-
[6]
C., and Chhatwal, J
Ayer, T., Zhang, C., Bonifonte, A., Spaulding, A. C., and Chhatwal, J. (2019). Prioritizing Hepatitis C Treatment in U . S . Prisons . Operations Research , 67(3):853--873
2019
-
[7]
Basole, R. C. and Bellamy, M. A. (2014). Supply Network Structure , Visibility , and Risk Diffusion : A Computational Approach . Decision Sciences , 45(4):753--789
2014
-
[8]
S., Mani, V., Venkatesh, V
Belhadi, A., Kamble, S. S., Mani, V., Venkatesh, V. G., and Shi, Y. (2021). Behavioral mechanisms influencing sustainable supply chain governance decision-making from a dyadic buyer-supplier perspective. International Journal of Production Economics , 236:108136
2021
Show all 103 references
-
[9]
K., Éltető, N., Griffiths, T
Binz, M., Akata, E., Bethge, M., Brändle, F., Callaway, F., Coda-Forno, J., Dayan, P., Demircan, C., Eckstein, M. K., Éltető, N., Griffiths, T. L., Haridi, S., Jagadish, A. K., Ji-An, L., Kipnis, A., Kumar, S., Ludwig, T., Mathony, M., Mattar, M., Modirshanechi, A., Nath, S. S...
2025
-
[10]
Bonabeau, E. (2002). Agent-based modeling: Methods and techniques for simulating human systems. Proceedings of the National Academy of Sciences , 99(suppl\_3):7280--7287
2002
-
[11]
Bowersox, D. J. and Closs, D. J. (1989). Simulation in Logistics : A Review of Present Practice and a Look to the Future . Journal of Business Logistics , 10(1):133--148
1989
-
[12]
Brito, R. P. and Miguel, P. L. S. (2017). Power, Governance , and Value in Collaboration : Differences between Buyer and Supplier Perspectives . Journal of Supply Chain Management , 53(2):61--87. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jscm.12134
2017 doi
-
[13]
Carbone, V., Rouquet, A., and Roussat, C. (2017). The Rise of Crowd Logistics : A New Way to Co - Create Logistics Value . Journal of Business Logistics , 38(4):238--252
2017
-
[14]
R., Rockwood, R
Carter, C. R., Rockwood, R. F., Patel, P. C., Bachrach, D., Bendoly, E., DuHadway, S., and Kaufmann, L. (2024). Experiments in supply chain management research: A systematic review and future directions. Journal of Business Logistics , 45(3):e12382
2024
-
[15]
E., Bell, J
Castillo, V. E., Bell, J. E., Mollenkopf, D. A., and Stank, T. P. (2021). Hybrid last mile delivery fleets with crowdsourcing: A systems view of managing the cost‐service trade‐off. Journal of Business Logistics , 43(1):36--61
2021
-
[16]
E., Bell, J
Castillo, V. E., Bell, J. E., Rose, W. J., and Rodrigues, A. M. (2018). Crowdsourcing last mile delivery: Strategic implications and future research directions. Journal of Business Logistics , 39(1):7--25
2018
-
[17]
E., Mollenkopf, D
Castillo, V. E., Mollenkopf, D. A., Bell, J. E., and Esper, T. L. (2022). Designing technology for on-demand delivery: The effect of customer tipping on crowdsourced driver behavior and last mile performance. Journal of Operations Management , 68(5):424--453. \_eprint: https:/...
2022 doi
-
[18]
J., and Benner, M
Chandrasekaran, A., Linderman, K., Sting, F. J., and Benner, M. J. (2016). Managing R & D Project Shifts in High - Tech Organizations : A Multi - Method Study . Production and Operations Management , 25(3):390--416
2016
-
[19]
L., Schonger, M., and Wickens, C
Chen, D. L., Schonger, M., and Wickens, C. (2016). oTree — An open-source platform for laboratory, online, and field experiments. Journal of Behavioral and Experimental Finance , 9:88--97
2016
-
[20]
Chen, K., Li, Y., and Linderman, K. (2022). Supply network resilience learning: An exploratory data analytics study. Decision Sciences , 53(1):8--27
2022
-
[21]
K., Chevalier, J
Chen, M. K., Chevalier, J. A., Rossi, P. E., and Oehlsen, E. (2019). The Value of Flexible Work : Evidence from Uber Drivers . Journal of Political Economy , 127(6):2735--2794. Publisher: The University of Chicago Press
2019
-
[22]
Y., Dooley, K
Choi, T. Y., Dooley, K. J., and Rungtusanatham, M. (2001). Supply networks and complex adaptive systems: control versus emergence. Journal of operations management , 19(3):351--366
2001
-
[23]
W., Cheng, L., and Ketchen Jr., D
Craighead, C. W., Cheng, L., and Ketchen Jr., D. J. (2024). Using middle-range theorizing to advance supply chain management research: A how-to primer and demonstration. Journal of Business Logistics , 45(3):e12381. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jbl.12381
2024 doi
-
[24]
Goldilocks
Craighead, C. W., Ketchen, D. J., and Cheng, L. (2016). “ Goldilocks ” Theorizing in Supply Chain Research : Balancing Scientific and Practical Utility via Middle - Range Theory . Transportation Journal , 55(3):241--257. Publisher: Penn State University Press
2016
-
[25]
and Savelsbergh, M
Dayarian, I. and Savelsbergh, M. (2020). Crowdshipping and Same ‐day Delivery : Employing In ‐store Customers to Deliver Online Orders . Production and Operations Management , 29(9):2153--2174
2020
-
[26]
R., and Kaufmann, L
Eckerd, S., DuHadway, S., Bendoly, E., Carter, C. R., and Kaufmann, L. (2021). On making experimental design choices: Discussions on the use and challenges of demand effects, incentives, deception, samples, and vignettes. Journal of Operations Management , 67(2):261--275. \_ep...
2021 doi
-
[27]
and Tibshirani, R
Efron, B. and Tibshirani, R. J. (1994). An introduction to the bootstrap . Chapman and Hall/CRC
1994
-
[28]
Evers, P. T. and Wan, X. (2012). Systems Analysis Using Simulation . Journal of Business Logistics , 33(2):80--89. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.0000-0000.2012.01041.x
2012
-
[29]
Gao, C., Lan, X., Li, N., Yuan, Y., Ding, J., Zhou, Z., Xu, F., and Li, Y. (2023a). Large Language Models Empowered Agent -based Modeling and Simulation : A Survey and Perspectives . arXiv:2312.11970 [cs]
-
[30]
Gao, Y., Li, M., and Sun, S. (2023b). Field experiments in operations management. Journal of Operations Management , 69(4):676--701. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/joom.1240
-
[31]
and Anna, I
Georgy, M. and Anna, I. (2020). semopy: A Python package for Structural Equation Modeling . Structural Equation Modeling: A Multidisciplinary Journal , 27(6):952--963. arXiv:1905.09376 [stat]
2020 arXiv
-
[32]
Ghaffarzadegan, N., Majumdar, A., Williams, R., and Hosseinichimeh, N. (2024). Generative agent‐based modeling: an introduction and tutorial. System Dynamics Review , 40(1):1--29
2024
-
[33]
Giannoccaro, I., Nair, A., and Choi, T. (2018). The Impact of Control and Complexity on Supply Network Performance : An Empirically Informed Investigation Using NK Simulation Analysis . Decision Sciences , 49(4):625--659. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.11...
2018 doi
-
[34]
Gibbons, D. E. and Samaddar, S. (2009). Designing Referral Network Structures and Decision Rules to Streamline Provision of Urgent Health and Human Services . Decision Sciences , 40(2):351--371
2009
-
[35]
Gurcan, O., Dikenelli, O., and Bernon, C. (2013). A generic testing framework for agent-based simulation models. Journal of Simulation , 7(3):183--201. Publisher: Taylor & Francis \_eprint: https://doi.org/10.1057/jos.2012.26
2013 doi
-
[36]
and Virrantaus, K
Hall, A. and Virrantaus, K. (2016). Visualizing the workings of agent-based models: Diagrams as a tool for communication and knowledge acquisition. Computers, Environment and Urban Systems , 58:1--11
2016
-
[37]
Hayes, A. F. (2017). Introduction to Mediation , Moderation , and Conditional Process Analysis , Second Edition : A Regression - Based Approach . Guilford Publications. Google-Books-ID: 6uk7DwAAQBAJ
2017
-
[38]
Junprung, E. (2023). Exploring the Intersection of Large Language Models and Agent - Based Modeling via Prompt Engineering . arXiv:2308.07411 [cs]
2023 arXiv
-
[39]
and Kelton, W
Kasaie, P. and Kelton, W. D. (2015). Guidelines for design and analysis in agent-based simulation studies. In 2015 Winter Simulation Conference ( WSC ) , pages 183--193. ISSN: 1558-4305
2015
-
[40]
A., Kashy, D
Kenny, D. A., Kashy, D. A., and Cook, W. L. (2006). Dyadic data analysis . Dyadic data analysis. Guilford Press, New York, NY, US. Pages: xix, 458
2006
-
[41]
Kleijnen, J. P. C. (1995). Verification and validation of simulation models. European Journal of Operational Research , 82(1):145--162
1995
-
[42]
Lakens, D. (2017). Equivalence Tests : A Practical Primer for t Tests , Correlations , and Meta - Analyses . Social Psychological and Personality Science , 8(4):355--362. Publisher: SAGE Publications Inc
2017
-
[43]
M., and Isager, P
Lakens, D., Scheel, A. M., and Isager, P. M. (2018). Equivalence Testing for Psychological Research : A Tutorial . Advances in Methods and Practices in Psychological Science , 1(2):259--269. Publisher: SAGE Publications Inc
2018
-
[44]
and Törnberg, P
Larooij, M. and Törnberg, P. (2025). Do Large Language Models Solve the Problems of Agent - Based Modeling ? A Critical Review of Generative Social Simulations
2025
-
[45]
Law, A. M. (2015). Simulation Modeling and Analysis . McGraw-Hill, New York, 5th edition
2015
-
[46]
Law, A. M. and Kelton, W. D. (1982). Simulation modeling and analysis . Mcgraw-hill New York
1982
-
[47]
S., Seo, Y
Lee, Y. S., Seo, Y. W., and Siemsen, E. (2018). Running Behavioral Operations Experiments Using Amazon 's Mechanical Turk . Production and Operations Management , 27(5):973--989. Publisher: SAGE Publications
2018
-
[48]
and Pun, H
Lei, Y. and Pun, H. (2023). Two- Sided Platform Competition in the Presence of Tip Baiting
2023
-
[49]
F., and Rahman, M
Li, M., Alam, Z., Bernardes, E., Giannoccaro, I., Skilton, P. F., and Rahman, M. S. (2021). Out of Sight , out of Mind ? Modeling the Impacts of Financial Squeeze on Extended Supply Chain Networks . Journal of Business Logistics , 42(2):233--263
2021
-
[50]
Liu, T., Yang, J., and Yin, Y. (2025). Toward LLM - Agent - Based Modeling of Transportation Systems : A Conceptual Framework . arXiv:2412.06681 [cs]
2025 arXiv
-
[51]
F., Zehnder, C., and Antonakis, J
Lonati, S., Quiroga, B. F., Zehnder, C., and Antonakis, J. (2018). On doing relevant and rigorous experiments: Review and recommendations. Journal of Operations Management , 64(1):19--40
2018
-
[52]
Lu, Y., Aleta, A., Du, C., Shi, L., and Moreno, Y. (2024). LLMs and generative agent-based models for complex systems research. Physics of Life Reviews , 51:283--293
2024
-
[53]
and Withiam, G
Lynn, M. and Withiam, G. (2008). Tipping and its alternatives: business considerations and directions for research. Journal of Services Marketing , 22(4):328--336
2008
-
[54]
M., and Harris, J
Lynn, M., Zinkhan, G. M., and Harris, J. (1993). Consumer Tipping : A Cross - Country Study . Journal of Consumer Research , 20(3):478
1993
-
[55]
Macal, C. M. (2016). Everything you need to know about agent-based modelling and simulation. Journal of Simulation , 10(2):144--156. Publisher: Taylor & Francis \_eprint: https://doi.org/10.1057/jos.2016.7
2016 doi
-
[56]
Masorgo, N., Mir, S., and Hofer, A. R. (2023). You're driving me crazy! How emotions elicited by negative driver behaviors impact customer outcomes in last mile delivery. Journal of Business Logistics , 44(4):666--692
2023
-
[57]
McGrath, J. E. (1981). Dilemmatics: The Study of Research Choices and Dilemmas . American Behavioral Scientist , 25(2):179--210. ERIC Number: EJ255600
1981
-
[58]
T., Gomes, R., and Krapfel, R
Mentzer, J. T., Gomes, R., and Krapfel, R. E. (1989). Physical distribution service: A fundamental marketing concept? Journal of the Academy of Marketing Science , 17(1):53--62
1989
-
[59]
Miao, W., Deng, Y., Wang, W., Liu, Y., and Tang, C. S. (2023). The effects of surge pricing on driver behavior in the ride-sharing market: Evidence from a quasi-experiment. Journal of Operations Management , 69(5):794--822. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10....
2023 doi
-
[60]
Miller, K. D. (2015). Agent- Based Modeling and Organization Studies : A critical realist perspective. Organization Studies , 36(2):175--196
2015
-
[61]
Mitsopoulos, K., Bose, R., Mather, B., Bhatia, A., Gluck, K., Dorr, B., Lebiere, C., and Pirolli, P. (2023). Psychologically- Valid Generative Agents : A Novel Approach to Agent - Based Modeling in Social Sciences . Proceedings of the AAAI Symposium Series , 2(1):340--348. Number: 1
2023
-
[62]
N., Lynch, D
Nyaga, G. N., Lynch, D. F., Marshall, D., and Ambrose, E. (2013). Power Asymmetry , Adaptation and Collaboration in Dyadic Relationships Involving a Powerful Partner . Journal of Supply Chain Management , 49(3):42--65. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/...
2013 doi
-
[63]
O'Brien, S. A. and Yurieff, K. (2020). People are Luring Instacart Shoppers with Big Tips -- and then changing them to zero
2020
-
[64]
and Schitter, C
Palan, S. and Schitter, C. (2018). Prolific.ac— A subject pool for online experiments. Journal of Behavioral and Experimental Finance , 17:22--27
2018
-
[65]
A., and Berry, L
Parasuraman, A., Zeithaml, V. A., and Berry, L. L. (1988). SERVQUAL : A Multiple - Item Scale for Measuring Consumer Perceptions of Service Quality . Journal of Retailing , 64(1):12--40. Publisher: Elsevier B.V
1988
-
[66]
S., O'Brien, J
Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023). Generative Agents : Interactive Simulacra of Human Behavior . arXiv:2304.03442 [cs]
2023 arXiv
-
[67]
Rainey, C. (2014). Arguing for a Negligible Effect . American Journal of Political Science , 58(4):1083--1091. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/ajps.12102
2014 doi
-
[68]
and Rust, R
Rand, W. and Rust, R. T. (2011). Agent-based modeling in marketing: Guidelines for rigor. International Journal of Research in Marketing , 28(3):181--193
2011
-
[69]
J., Griffis, S
Rao, S., Goldsby, T. J., Griffis, S. E., and Iyengar, D. (2011). Electronic Logistics Service Quality (e- LSQ ): Its Impact on the Customer ’s Purchase Satisfaction and Retention . Journal of Business Logistics , 32(2):167--179. \_eprint: https://onlinelibrary.wiley.com/doi/pd...
2011
-
[70]
and Grimm, C
Ribbink, D. and Grimm, C. M. (2014). The impact of cultural differences on buyer–supplier negotiations: An experimental study. Journal of Operations Management , 32(3):114--126
2014
-
[71]
Rungtusanatham, M., Wallin, C., and Eckerd, S. (2011). The Vignette in a Scenario - Based Role - Playing Experiment . Journal of Supply Chain Management , 47(3):9--16. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1745-493X.2011.03232.x
2011 arXiv
-
[72]
Rysman, M. (2009). The Economics of Two - Sided Markets . Journal of Economic Perspectives , 23(3):125--143
2009
-
[73]
Sargent, R. G. (2013). Verification and validation of simulation models. Journal of Simulation , 7(1):12--24. Publisher: Taylor & Francis \_eprint: https://doi.org/10.1057/jos.2012.20
2013 doi
-
[74]
W., Castillo, V
Saunders, L. W., Castillo, V. E., Rose, W. J., Dohmen, A. E., and Bell, J. E. (2025). Improving Driver Engagement in Delivery and Rideshare Services . Journal of Business Logistics , 46(2):e70011. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jbl.70011
2025 doi
-
[75]
Schelling, T. C. (1969). Models of Segregation . The American Economic Review , 59(2):488--493. Publisher: American Economic Association
1969
-
[76]
Shafer, S. M. and Smunt, T. L. (2004). Empirical simulation studies in operations management: context, trends, and research opportunities. Journal of Operations Management , 22(4):345--354
2004
-
[77]
R., Kumar, A., Shang, G., Thatcher, J., Fransoo, J
Shalpegin, T., Browning, T. R., Kumar, A., Shang, G., Thatcher, J., Fransoo, J. C., Holweg, M., and Lawson, B. (2025). Generative AI and Empirical Research Methods in Operations Management . Journal of Operations Management , n/a(n/a). \_eprint: https://onlinelibrary.wiley.com...
2025 doi
-
[78]
Shrout, P. E. and Bolger, N. (2002). Mediation in experimental and nonexperimental studies: New procedures and recommendations. Psychological Methods , 7(4):422--445. Publisher: American Psychological Association (APA)
2002
-
[79]
O., Macal, C
Siebers, P. O., Macal, C. M., Garnett, J., Buxton, D., and Pidd, M. (2010). Discrete-event simulation is dead, long live agent-based simulation! Journal of Simulation , 4(3):204--210
2010
-
[80]
Simchi-Levi, D., Mellou, K., Menache, I., and Pathuri, J. (2025). Large Language Models for Supply Chain Decisions . arXiv:2507.21502 [cs]
2025 arXiv
-
[81]
D., Marcucci, E., Gatta, V., and Claudel, C
Simoni, M. D., Marcucci, E., Gatta, V., and Claudel, C. G. (2020). Potential last-mile impacts of crowdshipping services: a simulation-based evaluation. Transportation , 47(4):1933--1954
2020
-
[82]
Smith, E. B. and Rand, W. (2018). Simulating Macro - Level Effects from Micro - Level Observations . Management Science , 64(11):5405--5421
2018
-
[83]
Spring, M., Faulconbridge, J., and Sarwar, A. (2022). How information technology automates and augments processes: Insights from Artificial - Intelligence -based systems in professional service operations. Journal of Operations Management , 68(6-7):592--618. \_eprint: https://...
2022 doi
-
[84]
P., Pellathy, D
Stank, T. P., Pellathy, D. A., In, J., Mollenkopf, D. A., and Bell, J. E. (2017). New Frontiers in Logistics Research : Theorizing at the Middle Range . Journal of Business Logistics , 38(1):6--17. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jbl.12151
2017 doi
-
[85]
L., and Hofer, A
Ta, H., Esper, T. L., and Hofer, A. R. (2018). Designing crowdsourced delivery systems: The effect of driver disclosure and ethnic similarity. Journal of Operations Management , 60(1):19--33
2018
-
[86]
L., Hofer, A
Ta, H., Esper, T. L., Hofer, A. R., and Sodero, A. C. (2025a). Reconceptualizing E - Logistics Service Quality ( E - LSQ ) in Emerging Contexts : The Case of Crowdsourced Delivery . Journal of Business Logistics , 46(1)
-
[87]
L., Rossiter Hofer, A., and Sodero, A
Ta, H., Esper, T. L., Rossiter Hofer, A., and Sodero, A. (2023). Crowdsourced delivery and customer assessments of e- Logistics Service Quality : An appraisal theory perspective. Journal of Business Logistics , 44(3):345--368
2023
-
[88]
R., Jin, Y
Ta, H., Hofer, A. R., Jin, Y. H., Peinkofer, S. T., and Sodero, A. (2025b). Designing scenario-based experiments in retail SCM : methodological approaches and practical insights. International Journal of Physical Distribution & Logistics Management , 55(1):94--117
-
[89]
Törnberg, A. (2019). Abstractions on steroids: A critical realist approach to computer simulations. Journal for the Theory of Social Behaviour , 49(1):127--143. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jtsb.12194
2019 doi
-
[90]
Vallat, R. (2018). Pingouin: statistics in Python . Journal of Open Source Software , 3(31):1026. Publisher: The Open Journal
2018
-
[91]
S., Agapiou, J
Vezhnevets, A. S., Agapiou, J. P., Aharon, A., Ziv, R., Matyas, J., Duéñez-Guzmán, E. A., Cunningham, W. A., Osindero, S., Karmon, D., and Leibo, J. Z. (2023). Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia . arXiv:2...
2023 arXiv
-
[92]
E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J....
2020
-
[93]
Wang, L., Rabinovich, E., and Richards, T. J. (2022). Scalability in Platforms for Local Groceries : An Examination of Indirect Network Economies . Production and Operations Management , 31(1):318--340. Publisher: SAGE Publications
2022
-
[94]
Wang, W., Yin, Y., and Xie, L. (2025). Effects of service encounter quality on courier and customer encounter satisfaction and loyalty in crowdsourcing logistics: an actor-partner interdependence model. Applied Economics , 57(31):4560--4575. Publisher: Routledge \_eprint: http...
2025
-
[95]
Warren, N., Hanson, S., and Yuan, H. (2021). Feeling Manipulated : How Tip Request Sequence Impacts Customers and Service Providers ? Journal of Service Research , 24(1):66--83. Publisher: SAGE Publications Inc
2021
-
[96]
Warren, N. B. and Hanson, S. (2025). Tipping privacy: The detrimental impact of observation on non-tip responses. Journal of Business Research , 186:115008
2025
-
[97]
Wu, Z., Peng, R., Han, X., Zheng, S., Zhang, Y., and Xiao, C. (2023). Smart Agent - Based Modeling : On the Use of Large Language Models in Computer Simulations . arXiv:2311.06330 [cs, econ, q-fin]
2023 arXiv
-
[98]
Xiao, B., Yin, Z., and Shan, Z. (2023). Simulating Public Administration Crisis : A Novel Generative Agent - Based Simulation System to Lower Technology Barriers in Social Science Research . arXiv:2311.06957 [cs]
2023 arXiv
-
[99]
Xie, C., Chen, C., Jia, F., Ye, Z., Lai, S., Shu, K., Gu, J., Bibi, A., Hu, Z., Jurgens, D., Evans, J., Torr, P., Ghanem, B., and Li, G. (2024). Can Large Language Model Agents Simulate Human Trust Behavior ? arXiv:2402.04559 [cs]
2024 arXiv
-
[100]
Yang, L., Choi, T.-M., and Shi, X. (2025). Risks in On - Demand Service Platform Operations : The Innovative Framework With SET Strategies . IEEE Transactions on Engineering Management , 72:2311--2329
2025
-
[101]
and Wang, H
Zhang, W. and Wang, H. (2024). To allow or not to allow tipping: tipping strategy choice for service platforms under competition. International Journal of Production Research , 62(23):8462--8484. Publisher: Taylor & Francis \_eprint: https://doi.org/10.1080/00207543.2024.2342578
2024
-
[102]
Zhao, K., Zuo, Z., and Blackhurst, J. V. (2019). Modelling supply chain adaptation for disruptions: An empirically grounded complex adaptive systems approach. Journal of Operations Management , 65(2):190--212
2019
-
[103]
K., Allen, B., Gretz, R
Zhou, Q. K., Allen, B., Gretz, R. T., and Houston, M. B. (2022). Platform Exploitation : When Service Agents Defect with Customers from Online Service Platforms . Journal of Marketing , 86(2):105--125. Publisher: SAGE Publications Inc
2022
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.