Pith. sign in

REVIEW 4 major objections 5 minor 103 references

Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A generative agent can pass behavior-equivalence tests yet follow artificial decision processes.

desk verdict A real dual-validation framework and a genuinely new LLM-vs-human comparison, but the process-fidelity score is a sign-blind significance counter, so the headline paradox needs a robustness check before it carries the weight the paper puts on it. read the letter →

arxiv 2508.20234 v1 pith:IICMEIL4 submitted 2025-08-27 cs.MA cs.AIcs.CY

classification cs.MAcs.AIcs.CY
keywords generativeagent-basedmodelsLLMvalidationequivalencetestingstructuralequationmodelingdyadicsatisfactionfooddeliveryplatformsmoderatedmediationbehavioralsimulation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that generative agent-based models (GABMs) for logistics and supply chain research must be validated at two levels: surface behavioral equivalence and decision-process fidelity. In a controlled food-delivery experiment comparing six LLMs with 957 human participants (477 dyads), the author shows that the two levels diverge: GPT-4o was equivalent on all three behavioral measures (tip change, joint satisfaction, differential satisfaction) yet matched only 5 of the 10 human decision pathways in a structural equation model, while GPT-4.1, Sonnet 3.5, Sonnet 4, and Mistral Medium 3 matched 8 of 10 pathways despite weaker surface equivalence. The claim is that choosing an LLM for simulation depends on the research question: aggregate outcomes may need only surface validation, but theory and policy questions need process validation. The paper offers a dual-validation framework that other LSCM dyads can adapt.

What carries the argument

The dual-validation framework is the paper's central mechanism. It pairs Two One-Sided Tests (TOST) for surface-level equivalence — testing whether each LLM's mean response on tip change, joint satisfaction, and differential satisfaction lies within a pre-specified margin of the human mean — with multi-group Structural Equation Modeling (SEM) for decision-process validation. The SEM specifies a moderated mediation model: service outcome affects tip change, tip adjustability moderates that path, tip visibility moderates the path from tip change to dyadic satisfaction, and both satisfaction outcomes are direct products of service outcome and tip change. Process fidelity is operationalized as t

What would settle it

Run the same dual validation in a different dyadic LSCM setting, such as buyer-supplier negotiation, where the human SEM is independently validated with think-aloud process data. If a model passes TOST equivalence on all outcome measures and also matches all ten human pathways, then the equivalence-versus-process paradox is not a general property of LLM agents.

Watch

Extended reading notes

Core claim

The paper's central claim is that surface-level behavioral equivalence does not guarantee that LLMs replicate human decision-making processes in GABM applications. This is demonstrated as an equivalence-versus-process paradox: some LLMs produce outcomes statistically indistinguishable from humans while employing artificial internal decision pathways, and models with lower outcome fidelity can have more human-like processes. The evidence comes from a 4×2×2 factorial experiment in food delivery, where six LLMs (480 dyads each) and 477 human dyads responded to identical vignettes. Surface equivalence was assessed with TOST tests at a ±0.2 SD margin; process fidelity was scored by how many of ei

Load-bearing premise

The entire process-fidelity comparison rests on the human structural equation model being the correct and complete account of decision making in this dyad; if that model is misspecified, the pathway-matching scores and the paradox they reveal are built on sand.

Editorial extensions

If this is right

  • Researchers running GABMs must declare which validation level their question demands; passing one level tells you nothing about the other.
  • LLM selection for operational simulations should be guided by the task: surface equivalence is enough for outcome prediction, while mechanism-oriented theory and policy analysis require process validation.
  • The 2025 model generations did not beat the 2024 generations on human behavioral equivalence in this context, so 'state of the art' status is not a proxy for behavioral fidelity.
  • If LLMs hold artificial decision pathways despite correct outputs, then simulated emergent phenomena — even when statistically plausible — may be attributed to the wrong causal mechanisms.
  • The study provides a template for validation in other dyadic LSCM contexts, with method menus for surface and process validation in each domain.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The equivalence-versus-process paradox may be a general property of LLM agents, not a quirk of food delivery: because model training optimizes text likelihood rather than psychological mechanism, outcome matching and process matching are decoupled incentives. Testing this in buyer-supplier or shipper-carrier dyads would show whether the paradox generalizes.
  • The process-fidelity score treats all ten pathways as equally important, so rankings could shift with a weighted scoring that emphasizes theoretically central paths (service outcome to satisfaction) over secondary ones; this is my own methodological suggestion, not the paper's.
  • The four models with 8/10 pathway fidelity each still deviated on two paths, so even the 'best' process match is partial; exact coefficient equality, not just significance matching, is a stricter bar the paper does not apply.
  • The finding that newer models were less surface-equivalent could reflect changes in instruction-following or response calibration rather than true behavioral drift; a longitudinal study with matched prompts would separate these.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a dual-validation framework for generative agent-based models (GABMs) in logistics and supply chain management: surface-level behavioral equivalence testing (TOST) and decision-process validation (SEM). Using a 4×2×2 vignette experiment on food-delivery customer–worker dyads, the author compares six LLMs against 957 human participants (477 dyads) and reports that GPT-4o passes surface-level equivalence on all three behavioral measures but has the lowest process fidelity (5/10 SEM pathway matches), while GPT-4.1, Sonnet 3.5, Sonnet 4, and Mistral Medium 3 match 8/10 human pathways despite weaker surface equivalence. The paper interprets this as an 'equivalence-versus-process paradox' and argues that both validation levels are necessary in GABM development.

Significance. The paper addresses a timely and important gap: systematic validation of LLM-based agents in behavioral LSCM research. The human baseline is large (477 dyads) and the experimental design is careful, including parallel vignettes and response formats for humans and LLMs. The proposed dual-validation idea is sensible and method-agnostic, and the study provides one of the first empirical demonstrations that surface-level output equivalence can diverge from decision-process fidelity. However, the central paradox rests on a process-fidelity metric that is a sign-blind, threshold-dependent count of significant pathways. Because the metric treats opposite-sign coefficients as matches and ignores effect sizes, the headline ranking and the claimed paradox are not as robust as the manuscript suggests. The human indirect effects are all non-significant, so part of the 'fidelity' score amounts to matching null patterns. These issues are fixable with additional analyses, but they are load-bearing for the main contribution.

major comments (4)
  1. [§4.5.4 and Table 5; Appendix C] The process-fidelity score counts a pathway as 'matched' if the LLM's p-value is on the same side of 0.05 as the human p-value, ignoring sign and magnitude. This is not a minor coding choice: the central paradox is operationalized through this metric. For example, the human Joint satisfaction ~ Tip change path is β=+0.051 (p<0.001, Table C.1), while GPT-4.1's path is β=-0.040 (p<0.001, Table C.3). Both are counted as significant matches, contributing to GPT-4.1's 8/10 fidelity score, even though the LLM reproduces the human path with the opposite sign. Similarly, Differential satisfaction ~ Service outcome has human β=0.936 (Table C.1) vs GPT-4.1 β=0.350 (Table C.3), yet both count as matches. A model that replicates human decisions through an opposite psychological mechanism can therefore be labeled 'high process fidelity.' I recommend re-scoring with sign agreement required and/or a st
  2. [§4.5.4, Table C.8] The human model shows non-significant indirect effects on both joint and differential satisfaction (p = 0.606 and p = 0.287). The process-fidelity metric therefore rewards LLMs for reproducing null indirect effects. But matching a null pattern is much weaker evidence of authentic process replication than matching a well-powered positive effect. The human sample (477 dyads) is not obviously underpowered for the direct effects, but the indirect effects are small and the bootstrap CIs are wide; the test may simply be insensitive. The paper should report the bootstrap power or equivalence bounds for these indirect effects, or at minimum explicitly acknowledge that the two indirect-effect 'matches' are matches to a null. As it stands, the process-fidelity ranking partly depends on LLMs correctly not showing something that the human data cannot strongly demonstrate.
  3. [§4.5.1 and Table 4] The TOST equivalence margin of ±0.2 SD is introduced without justification or sensitivity analysis. This margin is not merely a statistical detail: the surface-level ranking is the first pillar of the claimed paradox. For example, GPT-4o is the only model with equivalence on all three measures, but the p-values for tip change and differential satisfaction are close to 0.02 and 0.004, respectively. If a tighter margin (e.g., ±0.1 SD) were used, some of these equivalences might disappear; if a looser margin were used, more models would pass and the surface-level hierarchy would change. I request a rationale for ±0.2 SD (e.g., prior literature or practical relevance) and a sensitivity analysis over a plausible range of margins, reporting the resulting surface-level rankings and the implications for the paradox.
  4. [§4.5.4, multi-group SEM] The process-level validation compares each LLM's p-value pattern to the human p-value pattern, but it does not formally test whether the LLM coefficients differ from the human coefficients. A more rigorous multi-group approach would estimate a constrained model (e.g., fixing paths to the human estimates) and test the chi-square difference or use an equivalence test on individual coefficients. As written, the method treats a non-significant p-value in an LLM as a 'match' even when the coefficient estimate is far from the human estimate (e.g., Differential satisfaction ~ Service outcome: human 0.936 vs Mistral Medium 3 0.20, both significant, counted as a match). The paper should either implement formal tests of coefficient equality or clearly frame the process-fidelity score as a descriptive significance-pattern metric, not as evidence of process equivalence.
minor comments (5)
  1. [Table 4] The p-values in Table 4 are formatted inconsistently (e.g., 0.02, 0.00, 1.00, 0.34) and some are rounded to two decimals in a way that obscures exact values. Please use a consistent three-decimal format and mark values below 0.001 as '<0.001'.
  2. [§4.4.6] The 'white text haiku' quality-control measure deserves a clearer description. As written, it is not obvious how requesting a haiku in white text detects AI-generated responses. Clarify the mechanism or remove this sentence.
  3. [Table 3] Means and standard deviations in Table 3 are typeset with unusual spacing (e.g., '3 .61' instead of '3.61'). This appears to be a rendering issue, but it makes the table harder to read. Please fix the formatting.
  4. [§4.5.2] The sentence 'Three dyads were removed due to three human participants failing quality checks, making that participant's randomly assigned dyad unusable' is slightly confusing: it implies a 1:1 mapping from participant failure to dyad removal. Clarify how dyads were defined and how a single failed participant invalidated a dyad.
  5. [§5.3] The limitation about static vignettes is useful, but it could be expanded to note that the LLM agents do not interact with each other; they respond to a fixed scenario. This is relevant because the paper describes 'dyadic interactions' and 'emergent' behavior, but the current implementation is a sequential vignette response, not a multi-turn interaction. Please clarify the scope.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the validation compares independent LLM outputs against an external human SEM baseline; the sign-blind pathway metric is a validity concern, not a circular derivation.

full rationale

The paper's central claim is an empirical comparison, not a derivation. The human SEM (Table C.1) is estimated from 957 human participants; each LLM SEM (Tables C.2–C.7) is estimated from 480 dyads per model. The process-fidelity score counts significance matches across 8 direct paths and 2 indirect effects (Table 5). No parameter is fitted from LLM data to predict a human quantity; TOST equivalence margins (±0.2SD) are defined from human means, which is standard practice. The paper's self-citations (Castillo et al. 2018, 2021, 2022; Saunders et al. 2025) are domain literature and are not load-bearing for the validation framework or for the equivalence-versus-process claim. The sign-blindness of the pathway scoring (e.g., human Joint satisfaction ~ Tip change β=+0.051, p<0.001 vs GPT-4.1 β=−0.040, p<0.001, both counted as matches) is a substantive construct-validity limitation of the process-fidelity metric, but it does not make the comparison circular: the human baseline is independent, and the contradiction between surface equivalence and process fidelity is an observed empirical pattern, not a definitional identity. The paper's own limitation statements (static vignettes, aligned dyadic experiences) concern external validity, not circularity. Therefore no circular step is present.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new physical entities or theoretical constructs; the 'dual-validation framework' is a methodological procedure, not an invented entity. The load-bearing choices are parametric (equivalence margin, temperature, pathway list) and assumptions about the human SEM being the correct baseline. The most consequential is the equation-free but assumption-heavy process-fidelity metric, which counts pathway matches to a self-constructed human model. The structural-zero imputation for tip change is a data-modeling choice that could reasonably alter the moderation results.

free parameters (4)
  • TOST equivalence margin (±0.2 SD) = 0.2 SD
    Chosen without domain justification; the process fidelity rankings and the 'how many measures equivalent' scores depend directly on this margin. No sensitivity analysis is reported.
  • SEM pathway set (8 direct paths + 2 indirect effects) = 10 pathways
    The 'correct' human decision process is defined by the theory-driven model the authors chose. The LLM fidelity scores are counts of matches to this specific pathway list; a different theoretically justified model would change the rankings.
  • Bootstrap replications (5,000) = 5000
    Standard choice, not a major concern, but the significance classification of human indirect effects (all non-significant) is the anchor against which LLM artificial indirect effects are judged, so the inference threshold matters.
  • LLM temperature and top_p (0.7, 0.95) = temperature=0.7, top_p=0.95
    Hand-chosen to balance consistency and variation; equivalence and process results could shift at other sampling temperatures, and no robustness check is reported.
assumptions (5)
  • domain assumption The human SEM estimated from 477 dyads is the correct baseline representation of dyadic decision making in this context.
    The entire process-fidelity ranking is computed by matching LLM pathway significance patterns to this single human model (Section 4.5.4, Tables C.1-C.8). If the model is misspecified or underpowered, the rankings collapse.
  • domain assumption Surface-level equivalence is appropriately tested at the mean level with TOST over three outcome variables.
    The paper defines human equivalence as mean equivalence within ±0.2 SD on three measures. Distributional equivalence, variance equivalence, and response-pattern equivalence are not tested, so the 'surface-level' claim is narrower than it appears.
  • domain assumption LLM responses at temperature 0.7 with identical prompts are valid individual-level behavioral samples.
    The paper treats each of the 30 LLM replications per condition as an independent 'participant'. LLM responses within a model are not independent in the way human participants are; they share weights, training data, and prompt templates, which can deflate or inflate variance estimates (Section 4.4.5).
  • domain assumption The vignette-based paradigm captures the decision process of real dyadic food delivery encounters.
    The paper acknowledges in Limitations that responses are to static vignettes, not dynamic operational simulations. The validity of the whole framework for real platform behavior depends on vignette responses generalizing.
  • domain assumption Structural zeros imputed for tip change in non-adjustable conditions do not distort the SEM estimates.
    Imputing tip change = 0 for all non-adjustable dyads (Section 4.4.4) forces a large mass point at zero. The SEM treats this as continuous data, which can bias the moderation estimates that drive the tip adjustability findings.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research." pith.science (2026). https://pith.science/paper/IICMEIL4

@misc{pith2026250820234,
  author       = {Pith},
  title        = {Pith review of: Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IICMEIL4}},
  note         = {Machine review of arXiv:2508.20234}
}
read the original abstract

Generative Agent-Based Models (GABMs) powered by large language models (LLMs) offer promising potential for empirical logistics and supply chain management (LSCM) research by enabling realistic simulation of complex human behaviors. Unlike traditional agent-based models, GABMs generate human-like responses through natural language reasoning, which creates potential for new perspectives on emergent LSCM phenomena. However, the validity of LLMs as proxies for human behavior in LSCM simulations is unknown. This study evaluates LLM equivalence of human behavior through a controlled experiment examining dyadic customer-worker engagements in food delivery scenarios. I test six state-of-the-art LLMs against 957 human participants (477 dyads) using a moderated mediation design. This study reveals a need to validate GABMs on two levels: (1) human equivalence testing, and (2) decision process validation. Results reveal GABMs can effectively simulate human behaviors in LSCM; however, an equivalence-versus-process paradox emerges. While a series of Two One-Sided Tests (TOST) for equivalence reveals some LLMs demonstrate surface-level equivalence to humans, structural equation modeling (SEM) reveals artificial decision processes not present in human participants for some LLMs. These findings show GABMs as a potentially viable methodological instrument in LSCM with proper validation checks. The dual-validation framework also provides LSCM researchers with a guide to rigorous GABM development. For practitioners, this study offers evidence-based assessment for LLM selection for operational tasks.

Figures

Figures reproduced from arXiv: 2508.20234 by the authors.

Figure 1
Figure 1. Example of the worker-facing app when the customer has removed a tip after food delivery completion. [PITH_FULL_IMAGE:figures/full_fig_p012_1.png] view at source ↗
Figure 2
Figure 2. Hypothesized Moderated Mediation Model of Tipping Policy on Dyadic Satisfaction. [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. GABM Conceptual Model In the second stage, a worker receives notification of the delivery task. Depending on the experimental con￾dition (tip visibility), the worker may or may not see the customer’s initial tip amount when deciding whether 15 [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Overall Joint and Differential Dyadic Satisfaction Distributions [PITH_FULL_IMAGE:figures/full_fig_p021_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

103 extracted references · 67 canonical work pages

  1. [1]

    K., Lemarinier, P., and O’Hare, G

    Abar, S., Theodoropoulos, G. K., Lemarinier, P., and O’Hare, G. M. (2017). Agent Based Modelling and Simulation tools: A review of the state-of-art software. Computer Science Review , 24:13--33

  2. [2]

    and Meloy, M

    Abbey, J. and Meloy, M. (2017). Attention by design: Using attention checks to detect inattentive respondents and improve data quality. Journal of Operations Management , 53-56

  3. [3]

    C., and Sinchaisri, W

    Allon, G., Cohen, M. C., and Sinchaisri, W. P. (2023). The Impact of Behavioral and Economic Drivers on Gig Economy Workers . Manufacturing & Service Operations Management , 25(4):1376--1393. Publisher: INFORMS

  4. [4]

    and Pal, R

    Altay, N. and Pal, R. (2014). Information Diffusion among Agents : Implications for Humanitarian Operations . Production and Operations Management , 23(6):1015--1027

  5. [5]

    Ambra, T., Caris, A., and Macharis, C. (2021). Do You See What I See ? A Simulation Analysis of Order Bundling within a Transparent User Network in Geographic Space . Journal of Business Logistics , 42(1):167--190. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jbl.12237

  6. [6]

    C., and Chhatwal, J

    Ayer, T., Zhang, C., Bonifonte, A., Spaulding, A. C., and Chhatwal, J. (2019). Prioritizing Hepatitis C Treatment in U . S . Prisons . Operations Research , 67(3):853--873

  7. [7]

    Basole, R. C. and Bellamy, M. A. (2014). Supply Network Structure , Visibility , and Risk Diffusion : A Computational Approach . Decision Sciences , 45(4):753--789

  8. [8]

    S., Mani, V., Venkatesh, V

    Belhadi, A., Kamble, S. S., Mani, V., Venkatesh, V. G., and Shi, Y. (2021). Behavioral mechanisms influencing sustainable supply chain governance decision-making from a dyadic buyer-supplier perspective. International Journal of Production Economics , 236:108136

Show all 103 references
  1. [9]

    K., Éltető, N., Griffiths, T

    Binz, M., Akata, E., Bethge, M., Brändle, F., Callaway, F., Coda-Forno, J., Dayan, P., Demircan, C., Eckstein, M. K., Éltető, N., Griffiths, T. L., Haridi, S., Jagadish, A. K., Ji-An, L., Kipnis, A., Kumar, S., Ludwig, T., Mathony, M., Mattar, M., Modirshanechi, A., Nath, S. S...

  2. [10]

    Bonabeau, E. (2002). Agent-based modeling: Methods and techniques for simulating human systems. Proceedings of the National Academy of Sciences , 99(suppl\_3):7280--7287

  3. [11]

    Bowersox, D. J. and Closs, D. J. (1989). Simulation in Logistics : A Review of Present Practice and a Look to the Future . Journal of Business Logistics , 10(1):133--148

  4. [12]

    Brito, R. P. and Miguel, P. L. S. (2017). Power, Governance , and Value in Collaboration : Differences between Buyer and Supplier Perspectives . Journal of Supply Chain Management , 53(2):61--87. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jscm.12134

  5. [13]

    Carbone, V., Rouquet, A., and Roussat, C. (2017). The Rise of Crowd Logistics : A New Way to Co - Create Logistics Value . Journal of Business Logistics , 38(4):238--252

  6. [14]

    R., Rockwood, R

    Carter, C. R., Rockwood, R. F., Patel, P. C., Bachrach, D., Bendoly, E., DuHadway, S., and Kaufmann, L. (2024). Experiments in supply chain management research: A systematic review and future directions. Journal of Business Logistics , 45(3):e12382

  7. [15]

    E., Bell, J

    Castillo, V. E., Bell, J. E., Mollenkopf, D. A., and Stank, T. P. (2021). Hybrid last mile delivery fleets with crowdsourcing: A systems view of managing the cost‐service trade‐off. Journal of Business Logistics , 43(1):36--61

  8. [16]

    E., Bell, J

    Castillo, V. E., Bell, J. E., Rose, W. J., and Rodrigues, A. M. (2018). Crowdsourcing last mile delivery: Strategic implications and future research directions. Journal of Business Logistics , 39(1):7--25

  9. [17]

    E., Mollenkopf, D

    Castillo, V. E., Mollenkopf, D. A., Bell, J. E., and Esper, T. L. (2022). Designing technology for on-demand delivery: The effect of customer tipping on crowdsourced driver behavior and last mile performance. Journal of Operations Management , 68(5):424--453. \_eprint: https:/...

  10. [18]

    J., and Benner, M

    Chandrasekaran, A., Linderman, K., Sting, F. J., and Benner, M. J. (2016). Managing R & D Project Shifts in High - Tech Organizations : A Multi - Method Study . Production and Operations Management , 25(3):390--416

  11. [19]

    L., Schonger, M., and Wickens, C

    Chen, D. L., Schonger, M., and Wickens, C. (2016). oTree — An open-source platform for laboratory, online, and field experiments. Journal of Behavioral and Experimental Finance , 9:88--97

  12. [20]

    Chen, K., Li, Y., and Linderman, K. (2022). Supply network resilience learning: An exploratory data analytics study. Decision Sciences , 53(1):8--27

  13. [21]

    K., Chevalier, J

    Chen, M. K., Chevalier, J. A., Rossi, P. E., and Oehlsen, E. (2019). The Value of Flexible Work : Evidence from Uber Drivers . Journal of Political Economy , 127(6):2735--2794. Publisher: The University of Chicago Press

  14. [22]

    Y., Dooley, K

    Choi, T. Y., Dooley, K. J., and Rungtusanatham, M. (2001). Supply networks and complex adaptive systems: control versus emergence. Journal of operations management , 19(3):351--366

  15. [23]

    W., Cheng, L., and Ketchen Jr., D

    Craighead, C. W., Cheng, L., and Ketchen Jr., D. J. (2024). Using middle-range theorizing to advance supply chain management research: A how-to primer and demonstration. Journal of Business Logistics , 45(3):e12381. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jbl.12381

  16. [24]

    Goldilocks

    Craighead, C. W., Ketchen, D. J., and Cheng, L. (2016). “ Goldilocks ” Theorizing in Supply Chain Research : Balancing Scientific and Practical Utility via Middle - Range Theory . Transportation Journal , 55(3):241--257. Publisher: Penn State University Press

  17. [25]

    and Savelsbergh, M

    Dayarian, I. and Savelsbergh, M. (2020). Crowdshipping and Same ‐day Delivery : Employing In ‐store Customers to Deliver Online Orders . Production and Operations Management , 29(9):2153--2174

  18. [26]

    R., and Kaufmann, L

    Eckerd, S., DuHadway, S., Bendoly, E., Carter, C. R., and Kaufmann, L. (2021). On making experimental design choices: Discussions on the use and challenges of demand effects, incentives, deception, samples, and vignettes. Journal of Operations Management , 67(2):261--275. \_ep...

  19. [27]

    and Tibshirani, R

    Efron, B. and Tibshirani, R. J. (1994). An introduction to the bootstrap . Chapman and Hall/CRC

  20. [28]

    Evers, P. T. and Wan, X. (2012). Systems Analysis Using Simulation . Journal of Business Logistics , 33(2):80--89. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.0000-0000.2012.01041.x

  21. [29]

    Gao, C., Lan, X., Li, N., Yuan, Y., Ding, J., Zhou, Z., Xu, F., and Li, Y. (2023a). Large Language Models Empowered Agent -based Modeling and Simulation : A Survey and Perspectives . arXiv:2312.11970 [cs]

  22. [30]

    Gao, Y., Li, M., and Sun, S. (2023b). Field experiments in operations management. Journal of Operations Management , 69(4):676--701. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/joom.1240

  23. [31]

    and Anna, I

    Georgy, M. and Anna, I. (2020). semopy: A Python package for Structural Equation Modeling . Structural Equation Modeling: A Multidisciplinary Journal , 27(6):952--963. arXiv:1905.09376 [stat]

  24. [32]

    Ghaffarzadegan, N., Majumdar, A., Williams, R., and Hosseinichimeh, N. (2024). Generative agent‐based modeling: an introduction and tutorial. System Dynamics Review , 40(1):1--29

  25. [33]

    Giannoccaro, I., Nair, A., and Choi, T. (2018). The Impact of Control and Complexity on Supply Network Performance : An Empirically Informed Investigation Using NK Simulation Analysis . Decision Sciences , 49(4):625--659. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.11...

  26. [34]

    Gibbons, D. E. and Samaddar, S. (2009). Designing Referral Network Structures and Decision Rules to Streamline Provision of Urgent Health and Human Services . Decision Sciences , 40(2):351--371

  27. [35]

    Gurcan, O., Dikenelli, O., and Bernon, C. (2013). A generic testing framework for agent-based simulation models. Journal of Simulation , 7(3):183--201. Publisher: Taylor & Francis \_eprint: https://doi.org/10.1057/jos.2012.26

  28. [36]

    and Virrantaus, K

    Hall, A. and Virrantaus, K. (2016). Visualizing the workings of agent-based models: Diagrams as a tool for communication and knowledge acquisition. Computers, Environment and Urban Systems , 58:1--11

  29. [37]

    Hayes, A. F. (2017). Introduction to Mediation , Moderation , and Conditional Process Analysis , Second Edition : A Regression - Based Approach . Guilford Publications. Google-Books-ID: 6uk7DwAAQBAJ

  30. [38]

    Junprung, E. (2023). Exploring the Intersection of Large Language Models and Agent - Based Modeling via Prompt Engineering . arXiv:2308.07411 [cs]

  31. [39]

    and Kelton, W

    Kasaie, P. and Kelton, W. D. (2015). Guidelines for design and analysis in agent-based simulation studies. In 2015 Winter Simulation Conference ( WSC ) , pages 183--193. ISSN: 1558-4305

  32. [40]

    A., Kashy, D

    Kenny, D. A., Kashy, D. A., and Cook, W. L. (2006). Dyadic data analysis . Dyadic data analysis. Guilford Press, New York, NY, US. Pages: xix, 458

  33. [41]

    Kleijnen, J. P. C. (1995). Verification and validation of simulation models. European Journal of Operational Research , 82(1):145--162

  34. [42]

    Lakens, D. (2017). Equivalence Tests : A Practical Primer for t Tests , Correlations , and Meta - Analyses . Social Psychological and Personality Science , 8(4):355--362. Publisher: SAGE Publications Inc

  35. [43]

    M., and Isager, P

    Lakens, D., Scheel, A. M., and Isager, P. M. (2018). Equivalence Testing for Psychological Research : A Tutorial . Advances in Methods and Practices in Psychological Science , 1(2):259--269. Publisher: SAGE Publications Inc

  36. [44]

    and Törnberg, P

    Larooij, M. and Törnberg, P. (2025). Do Large Language Models Solve the Problems of Agent - Based Modeling ? A Critical Review of Generative Social Simulations

  37. [45]

    Law, A. M. (2015). Simulation Modeling and Analysis . McGraw-Hill, New York, 5th edition

  38. [46]

    Law, A. M. and Kelton, W. D. (1982). Simulation modeling and analysis . Mcgraw-hill New York

  39. [47]

    S., Seo, Y

    Lee, Y. S., Seo, Y. W., and Siemsen, E. (2018). Running Behavioral Operations Experiments Using Amazon 's Mechanical Turk . Production and Operations Management , 27(5):973--989. Publisher: SAGE Publications

  40. [48]

    and Pun, H

    Lei, Y. and Pun, H. (2023). Two- Sided Platform Competition in the Presence of Tip Baiting

  41. [49]

    F., and Rahman, M

    Li, M., Alam, Z., Bernardes, E., Giannoccaro, I., Skilton, P. F., and Rahman, M. S. (2021). Out of Sight , out of Mind ? Modeling the Impacts of Financial Squeeze on Extended Supply Chain Networks . Journal of Business Logistics , 42(2):233--263

  42. [50]

    Liu, T., Yang, J., and Yin, Y. (2025). Toward LLM - Agent - Based Modeling of Transportation Systems : A Conceptual Framework . arXiv:2412.06681 [cs]

  43. [51]

    F., Zehnder, C., and Antonakis, J

    Lonati, S., Quiroga, B. F., Zehnder, C., and Antonakis, J. (2018). On doing relevant and rigorous experiments: Review and recommendations. Journal of Operations Management , 64(1):19--40

  44. [52]

    Lu, Y., Aleta, A., Du, C., Shi, L., and Moreno, Y. (2024). LLMs and generative agent-based models for complex systems research. Physics of Life Reviews , 51:283--293

  45. [53]

    and Withiam, G

    Lynn, M. and Withiam, G. (2008). Tipping and its alternatives: business considerations and directions for research. Journal of Services Marketing , 22(4):328--336

  46. [54]

    M., and Harris, J

    Lynn, M., Zinkhan, G. M., and Harris, J. (1993). Consumer Tipping : A Cross - Country Study . Journal of Consumer Research , 20(3):478

  47. [55]

    Macal, C. M. (2016). Everything you need to know about agent-based modelling and simulation. Journal of Simulation , 10(2):144--156. Publisher: Taylor & Francis \_eprint: https://doi.org/10.1057/jos.2016.7

  48. [56]

    Masorgo, N., Mir, S., and Hofer, A. R. (2023). You're driving me crazy! How emotions elicited by negative driver behaviors impact customer outcomes in last mile delivery. Journal of Business Logistics , 44(4):666--692

  49. [57]

    McGrath, J. E. (1981). Dilemmatics: The Study of Research Choices and Dilemmas . American Behavioral Scientist , 25(2):179--210. ERIC Number: EJ255600

  50. [58]

    T., Gomes, R., and Krapfel, R

    Mentzer, J. T., Gomes, R., and Krapfel, R. E. (1989). Physical distribution service: A fundamental marketing concept? Journal of the Academy of Marketing Science , 17(1):53--62

  51. [59]

    Miao, W., Deng, Y., Wang, W., Liu, Y., and Tang, C. S. (2023). The effects of surge pricing on driver behavior in the ride-sharing market: Evidence from a quasi-experiment. Journal of Operations Management , 69(5):794--822. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10....

  52. [60]

    Miller, K. D. (2015). Agent- Based Modeling and Organization Studies : A critical realist perspective. Organization Studies , 36(2):175--196

  53. [61]

    Mitsopoulos, K., Bose, R., Mather, B., Bhatia, A., Gluck, K., Dorr, B., Lebiere, C., and Pirolli, P. (2023). Psychologically- Valid Generative Agents : A Novel Approach to Agent - Based Modeling in Social Sciences . Proceedings of the AAAI Symposium Series , 2(1):340--348. Number: 1

  54. [62]

    N., Lynch, D

    Nyaga, G. N., Lynch, D. F., Marshall, D., and Ambrose, E. (2013). Power Asymmetry , Adaptation and Collaboration in Dyadic Relationships Involving a Powerful Partner . Journal of Supply Chain Management , 49(3):42--65. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/...

  55. [63]

    O'Brien, S. A. and Yurieff, K. (2020). People are Luring Instacart Shoppers with Big Tips -- and then changing them to zero

  56. [64]

    and Schitter, C

    Palan, S. and Schitter, C. (2018). Prolific.ac— A subject pool for online experiments. Journal of Behavioral and Experimental Finance , 17:22--27

  57. [65]

    A., and Berry, L

    Parasuraman, A., Zeithaml, V. A., and Berry, L. L. (1988). SERVQUAL : A Multiple - Item Scale for Measuring Consumer Perceptions of Service Quality . Journal of Retailing , 64(1):12--40. Publisher: Elsevier B.V

  58. [66]

    S., O'Brien, J

    Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023). Generative Agents : Interactive Simulacra of Human Behavior . arXiv:2304.03442 [cs]

  59. [67]

    Rainey, C. (2014). Arguing for a Negligible Effect . American Journal of Political Science , 58(4):1083--1091. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/ajps.12102

  60. [68]

    and Rust, R

    Rand, W. and Rust, R. T. (2011). Agent-based modeling in marketing: Guidelines for rigor. International Journal of Research in Marketing , 28(3):181--193

  61. [69]

    J., Griffis, S

    Rao, S., Goldsby, T. J., Griffis, S. E., and Iyengar, D. (2011). Electronic Logistics Service Quality (e- LSQ ): Its Impact on the Customer ’s Purchase Satisfaction and Retention . Journal of Business Logistics , 32(2):167--179. \_eprint: https://onlinelibrary.wiley.com/doi/pd...

  62. [70]

    and Grimm, C

    Ribbink, D. and Grimm, C. M. (2014). The impact of cultural differences on buyer–supplier negotiations: An experimental study. Journal of Operations Management , 32(3):114--126

  63. [71]

    Rungtusanatham, M., Wallin, C., and Eckerd, S. (2011). The Vignette in a Scenario - Based Role - Playing Experiment . Journal of Supply Chain Management , 47(3):9--16. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1745-493X.2011.03232.x

  64. [72]

    Rysman, M. (2009). The Economics of Two - Sided Markets . Journal of Economic Perspectives , 23(3):125--143

  65. [73]

    Sargent, R. G. (2013). Verification and validation of simulation models. Journal of Simulation , 7(1):12--24. Publisher: Taylor & Francis \_eprint: https://doi.org/10.1057/jos.2012.20

  66. [74]

    W., Castillo, V

    Saunders, L. W., Castillo, V. E., Rose, W. J., Dohmen, A. E., and Bell, J. E. (2025). Improving Driver Engagement in Delivery and Rideshare Services . Journal of Business Logistics , 46(2):e70011. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jbl.70011

  67. [75]

    Schelling, T. C. (1969). Models of Segregation . The American Economic Review , 59(2):488--493. Publisher: American Economic Association

  68. [76]

    Shafer, S. M. and Smunt, T. L. (2004). Empirical simulation studies in operations management: context, trends, and research opportunities. Journal of Operations Management , 22(4):345--354

  69. [77]

    R., Kumar, A., Shang, G., Thatcher, J., Fransoo, J

    Shalpegin, T., Browning, T. R., Kumar, A., Shang, G., Thatcher, J., Fransoo, J. C., Holweg, M., and Lawson, B. (2025). Generative AI and Empirical Research Methods in Operations Management . Journal of Operations Management , n/a(n/a). \_eprint: https://onlinelibrary.wiley.com...

  70. [78]

    Shrout, P. E. and Bolger, N. (2002). Mediation in experimental and nonexperimental studies: New procedures and recommendations. Psychological Methods , 7(4):422--445. Publisher: American Psychological Association (APA)

  71. [79]

    O., Macal, C

    Siebers, P. O., Macal, C. M., Garnett, J., Buxton, D., and Pidd, M. (2010). Discrete-event simulation is dead, long live agent-based simulation! Journal of Simulation , 4(3):204--210

  72. [80]

    Simchi-Levi, D., Mellou, K., Menache, I., and Pathuri, J. (2025). Large Language Models for Supply Chain Decisions . arXiv:2507.21502 [cs]

  73. [81]

    D., Marcucci, E., Gatta, V., and Claudel, C

    Simoni, M. D., Marcucci, E., Gatta, V., and Claudel, C. G. (2020). Potential last-mile impacts of crowdshipping services: a simulation-based evaluation. Transportation , 47(4):1933--1954

  74. [82]

    Smith, E. B. and Rand, W. (2018). Simulating Macro - Level Effects from Micro - Level Observations . Management Science , 64(11):5405--5421

  75. [83]

    Spring, M., Faulconbridge, J., and Sarwar, A. (2022). How information technology automates and augments processes: Insights from Artificial - Intelligence -based systems in professional service operations. Journal of Operations Management , 68(6-7):592--618. \_eprint: https://...

  76. [84]

    P., Pellathy, D

    Stank, T. P., Pellathy, D. A., In, J., Mollenkopf, D. A., and Bell, J. E. (2017). New Frontiers in Logistics Research : Theorizing at the Middle Range . Journal of Business Logistics , 38(1):6--17. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jbl.12151

  77. [85]

    L., and Hofer, A

    Ta, H., Esper, T. L., and Hofer, A. R. (2018). Designing crowdsourced delivery systems: The effect of driver disclosure and ethnic similarity. Journal of Operations Management , 60(1):19--33

  78. [86]

    L., Hofer, A

    Ta, H., Esper, T. L., Hofer, A. R., and Sodero, A. C. (2025a). Reconceptualizing E - Logistics Service Quality ( E - LSQ ) in Emerging Contexts : The Case of Crowdsourced Delivery . Journal of Business Logistics , 46(1)

  79. [87]

    L., Rossiter Hofer, A., and Sodero, A

    Ta, H., Esper, T. L., Rossiter Hofer, A., and Sodero, A. (2023). Crowdsourced delivery and customer assessments of e- Logistics Service Quality : An appraisal theory perspective. Journal of Business Logistics , 44(3):345--368

  80. [88]

    R., Jin, Y

    Ta, H., Hofer, A. R., Jin, Y. H., Peinkofer, S. T., and Sodero, A. (2025b). Designing scenario-based experiments in retail SCM : methodological approaches and practical insights. International Journal of Physical Distribution & Logistics Management , 55(1):94--117

  81. [89]

    Törnberg, A. (2019). Abstractions on steroids: A critical realist approach to computer simulations. Journal for the Theory of Social Behaviour , 49(1):127--143. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jtsb.12194

  82. [90]

    Vallat, R. (2018). Pingouin: statistics in Python . Journal of Open Source Software , 3(31):1026. Publisher: The Open Journal

  83. [91]

    S., Agapiou, J

    Vezhnevets, A. S., Agapiou, J. P., Aharon, A., Ziv, R., Matyas, J., Duéñez-Guzmán, E. A., Cunningham, W. A., Osindero, S., Karmon, D., and Leibo, J. Z. (2023). Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia . arXiv:2...

  84. [92]

    E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S

    Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J....

  85. [93]

    Wang, L., Rabinovich, E., and Richards, T. J. (2022). Scalability in Platforms for Local Groceries : An Examination of Indirect Network Economies . Production and Operations Management , 31(1):318--340. Publisher: SAGE Publications

  86. [94]

    Wang, W., Yin, Y., and Xie, L. (2025). Effects of service encounter quality on courier and customer encounter satisfaction and loyalty in crowdsourcing logistics: an actor-partner interdependence model. Applied Economics , 57(31):4560--4575. Publisher: Routledge \_eprint: http...

  87. [95]

    Warren, N., Hanson, S., and Yuan, H. (2021). Feeling Manipulated : How Tip Request Sequence Impacts Customers and Service Providers ? Journal of Service Research , 24(1):66--83. Publisher: SAGE Publications Inc

  88. [96]

    Warren, N. B. and Hanson, S. (2025). Tipping privacy: The detrimental impact of observation on non-tip responses. Journal of Business Research , 186:115008

  89. [97]

    Wu, Z., Peng, R., Han, X., Zheng, S., Zhang, Y., and Xiao, C. (2023). Smart Agent - Based Modeling : On the Use of Large Language Models in Computer Simulations . arXiv:2311.06330 [cs, econ, q-fin]

  90. [98]

    Xiao, B., Yin, Z., and Shan, Z. (2023). Simulating Public Administration Crisis : A Novel Generative Agent - Based Simulation System to Lower Technology Barriers in Social Science Research . arXiv:2311.06957 [cs]

  91. [99]

    Xie, C., Chen, C., Jia, F., Ye, Z., Lai, S., Shu, K., Gu, J., Bibi, A., Hu, Z., Jurgens, D., Evans, J., Torr, P., Ghanem, B., and Li, G. (2024). Can Large Language Model Agents Simulate Human Trust Behavior ? arXiv:2402.04559 [cs]

  92. [100]

    Yang, L., Choi, T.-M., and Shi, X. (2025). Risks in On - Demand Service Platform Operations : The Innovative Framework With SET Strategies . IEEE Transactions on Engineering Management , 72:2311--2329

  93. [101]

    and Wang, H

    Zhang, W. and Wang, H. (2024). To allow or not to allow tipping: tipping strategy choice for service platforms under competition. International Journal of Production Research , 62(23):8462--8484. Publisher: Taylor & Francis \_eprint: https://doi.org/10.1080/00207543.2024.2342578

  94. [102]

    Zhao, K., Zuo, Z., and Blackhurst, J. V. (2019). Modelling supply chain adaptation for disruptions: An empirically grounded complex adaptive systems approach. Journal of Operations Management , 65(2):190--212

  95. [103]

    K., Allen, B., Gretz, R

    Zhou, Q. K., Allen, B., Gretz, R. T., and Houston, M. B. (2022). Platform Exploitation : When Service Agents Defect with Customers from Online Service Platforms . Journal of Marketing , 86(2):105--125. Publisher: SAGE Publications Inc

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.