{"id":"2c124972-d74f-477c-82f8-c5e80d8c7f53","arxiv_id":"2508.11416","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AIM-Bench shows LLM inventory agents exhibit human-like pull-to-center and bullwhip biases, partly mitigated by cognitive reflection and information sharing.","lead":"AIM-Bench is a new benchmark that tests AI agents as inventory managers under uncertain demand and scores their decision biases. It finds large language models show human-like biases, such as pull-to-center and bullwhip effects, and that certain interventions can reduce them.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract claims LLM biases are 'similar to humans' but does not state that human controls ran the identical task; with the supplied full text corrupted, the human-comparison basis is unverified.","rationale":"The reader's weakest assumption concerned framing artifacts versus intrinsic LLM biases. My concern is more specific: the 'human-like' conclusion requires a valid human comparison, and the abstract does not establish that one exists. Because the full text is corrupted in the provided copy, the protocol cannot be inspected, so the correct disposition remains UNVERDICTED rather than ACCEPT or REJECT. I do not see an internal inconsistency in the abstract itself, but the evidence needed to support the core comparative claim is absent from the available material. If the clean text reveals a proper human control group with matched task conditions, the concern would be resolved. If not, the central claim would be substantially weakened. Thus my read does not change the reader's verdict; it sharpens the reason for remaining unverified.","tokens_in":8824,"tokens_out":2713,"duration_ms":37126,"concrete_test":"Obtain the clean manuscript or arXiv source and locate the experimental protocol. Check whether human participants were recruited and run through the exact same AIM-Bench interface (same instructions, demand histories, unit scales, and order-input format). If yes, recompute the pull-to-center and bullwhip metrics on the human data and test whether LLM and human distributions overlap after normalizing by mean demand. If no human data exists or the human comparison comes only from prior literature, the 'similar to humans' claim is not supported by AIM-Bench. As an additional framing check, rerun one condition with an altered numerical scale (e.g., demands expressed in hundreds instead of units) and see whether bias magnitudes shift materially.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference is that LLM agents exhibit decision biases 'similar to those observed in human beings' and that cognitive reflection and information sharing mitigate them. For this claim to hold, the same bias metrics must be measured in LLMs and humans under comparable conditions. The abstract never states that a human control group was run through the same AIM-Bench task, with the same prompt framing, demand sequences, and response scale. If the 'human-like' comparison is instead qualitative alignment with published behavioral experiments that used different cover stories, number scales, and demand distributions, the observed pull-to-center and bullwhip metrics could be artifacts of the LLM's surface-level response tendencies (e.g., averaging or scale anchoring) rather than evidence of shared decision biases. The supplied full text is corrupted/unreadable, so the paper as available provides no protocol details, no explicit human-participant section, and no statistical comparison against human data. This is not a claim that the authors are wrong; it is a missing link between the measured LLM ordering behavior and the paper's headline interpretation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AIM-Bench, a benchmark designed to evaluate the decision-making behavior of LLM-based agents acting as inventory managers in uncertain supply chain scenarios. According to the abstract, the benchmark runs a diverse series of inventory replenishment experiments, and the results reveal that different LLMs exhibit decision biases, including pull-to-center and bullwhip effects, that are 'similar to those observed in human beings.' The abstract further claims that two interventions—cognitive reflection and information sharing—can mitigate these biases. The paper positions this as evidence that LLM agents may inherit human-like decision biases in operations contexts and argues for caution when deploying them in inventory management, as well as for developing human-centered decision support systems. The full text provided to the referee is heavily corrupted and largely unreadable, so the assessment below rests primarily on the abstract and the fragments that can be recovered.","tokens_in":9103,"tokens_out":3561,"duration_ms":45291,"significance":"If the claims are correct, the paper addresses a timely and important question: whether LLM agents used in supply chain and inventory management inherit systematic human biases and whether those biases can be reduced by prompt-level or information-level interventions. The idea of a benchmark specifically targeting decision biases in LLM inventory managers is valuable, and the mitigation experiments, if properly designed, could have practical implications for human-AI collaboration in operations. However, the available material does not provide enough detail to verify the central measurements, the human comparison, or the mitigation results. No machine-checked proofs, reproducible code, or full protocol are visible in the supplied text. The significance is therefore conditional: the contribution is potentially strong, but the evidence as presented is insufficient for a soundness assessment.","major_comments":[{"comment":"The headline claim is that LLMs exhibit decision biases 'similar to those observed in human beings.' The abstract gives no indication that human participants completed the same AIM-Bench task under the same prompt framing, demand sequences, and response scale. If the 'human-like' comparison is instead qualitative alignment with published behavioral experiments that used different cover stories and scales, the measured pull-to-center and bullwhip indices may reflect prompt-induced averaging or scale anchoring rather than a shared decision bias. The authors must identify the human comparison data, show that humans performed the identical task, and report direct statistical comparisons (bias indices, confidence intervals, and tests) for each LLM relative to the human data.","section":"Abstract / human comparison"},{"comment":"The readable portion of the manuscript contains no methodological detail: LLM model versions, prompt templates, temperature and decoding settings, number of independent runs or episodes, demand distributions, cost and lead-time parameters, and the exact definitions of the pull-to-center and bullwhip metrics are all absent. Without this information, the measured biases cannot be reproduced and could be one-off stochastic artifacts. At minimum, the paper must report these details and include variance or error bars for every reported bias index.","section":"Methods / protocol"},{"comment":"The abstract states that cognitive reflection and information sharing mitigate the pull-to-center and bullwhip effects, but the readable material does not define or quantify either intervention. The reader cannot determine how cognitive reflection was implemented (e.g., chain-of-thought prompting, self-consistency, or another manipulation), what information was shared, or whether the reported reductions are statistically significant and consistent across models. Because the mitigation results are a central contribution, they need full operationalization, effect sizes, and significance tests.","section":"Results / mitigation claims"},{"comment":"The supplied full text is largely unreadable due to encoding corruption; no section, table, or equation can currently be checked. While this may be an artifact of the review pipeline rather than the authors' fault, the manuscript as provided to the referee does not contain a verifiable protocol, results, or references. As it stands, a full evaluation is impossible, and the authors should be asked to provide a clean, readable version.","section":"Full text / legibility"}],"minor_comments":[{"comment":"The phrase 'pave the way for mitigating human decision bias' overstates the scope, since the study evaluates LLM agents, not human decision makers. Suggest rewording to focus on supporting human decision-makers through better LLM agents.","section":"Abstract, final sentence"},{"comment":"The abstract mentions 'e.g. framing effect, etc.' but the subsequent summary of findings only mentions pull-to-center and bullwhip effects. Either define the full set of biases evaluated in the benchmark or align the abstract with the actual contributions.","section":"Abstract, bias terminology"},{"comment":"If human data are not collected within the same task, the phrase 'similar to those observed in human beings' should be qualified as 'qualitatively consistent with published behavioral results,' with a clear statement of the limitations of such a comparison.","section":"General"}],"recommendation":"uncertain","confidential_remarks":"The central issue is that the full text is unreadable, so the referee cannot verify any of the paper's load-bearing claims. Before any further review, the editor should request a clean PDF or source version. The paper's topic and proposed benchmark are potentially suitable for an AI/operations audience, but the current submission is not reviewable in its provided form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: AIM-Bench is a reasonable idea—putting LLM agents into inventory replenishment tasks and measuring known behavioral biases (pull-to-center, bullwhip) plus mitigation attempts. That is a useful niche. But as far as I can tell from the evidence here, the paper is a claim, not a demonstration. The abstract reports results but no method; the full text I have is corrupted, so I can't check the actual experiments.\n\nCredit where it is due: the benchmark concept is new in the LLM-agent evaluation space, and the mitigation angle (cognitive reflection, information sharing) is a practical extension. The authors are aware of the relevant literature. The bias selection is sensible for inventory management. If the implementation is sound, this would be a solid contribution to AI operations.\n\nSoft spots: first, no explicit human control group is mentioned in the abstract. The claim that LLM biases are 'similar to those observed in human beings' requires that humans ran the same task with the same prompts and response scale; otherwise it is just a qualitative analogy. The stress-test note is right to flag this. Second, there are no details on model versions, number of trials, prompt templates, or statistical tests. That is a serious verification gap. Third, the text I was sent is unreadable due to encoding corruption, so I could not evaluate the protocol. If the actual manuscript contains these details, my concerns may be moot.\n\nWho is this for: researchers building LLM agents for supply chain or operations decisions, and people studying human-like biases in LLMs. It could be a useful reference if the evidence holds up.\n\nMy recommendation: I would not cite it yet, but I would not desk-reject it either. The idea deserves a serious referee. If you can obtain a readable version, send it out and ask reviewers specifically to verify whether the 'human-like' comparison uses actual human participants and whether the bias metrics are computed correctly. If the authors have the data, this could become a benchmark the field actually uses.","headline":"Plausible benchmark for inventory-decision biases in LLMs, but the abstract alone can't support the 'human-like' claim and the full text is unreadable.","tokens_in":9482,"tokens_out":2151,"would_cite":false,"duration_ms":25921,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that LLM agents acting as inventory managers reproduce two human decision biases — pull-to-center and the bullwhip effect — and that both can be reduced by cognitive reflection prompts and information sharing.","keywords":["LLM agents","inventory management","decision bias","pull-to-center effect","bullwhip effect","supply chain uncertainty","AIM-Bench","cognitive reflection"],"falsifier":"Re-run the same replenishment experiments with the same expected costs but with demand expressed in different units or scales; if the pull-to-center effect shifts substantially with the frame, the claimed human-like bias is partly a prompt artifact, not a stable property of the models.","tokens_in":8809,"feed_emoji":"📦","tokens_out":2740,"duration_ms":29697,"temperature":0.7,"pith_summary":"This paper introduces AIM-Bench, a benchmark built from inventory replenishment experiments under uncertain demand. It reports that several large language models, when acting as inventory managers, show two biases long observed in human decision-makers: pull-to-center (ordering closer to average demand than optimal) and the bullwhip effect (amplifying demand variability up the supply chain). The paper also reports that two interventions temper these biases: prompting agents to reflect cognitively, and sharing demand information across the chain. If these findings hold, AI agents deployed in supply-chain operations will inherit behavioural biases that can be partially corrected by how they are prompted and what information they see.","feed_headline":"LLM inventory agents show human-like ordering bias","feed_subtitle":"AIM-Bench measures pull-to-center and bullwhip in AI supply decisions; reflection prompts and shared demand data shrink both.","key_machinery":"AIM-Bench itself: a benchmark of inventory replenishment experiments in which an LLM agent repeatedly chooses order quantities under stochastic demand, either alone or in a multi-stage chain. The benchmark quantifies two named biases — pull-to-center, the tendency to order closer to the mean demand than cost-optimal, and bullwhip, the amplification of order variability upstream — and provides intervention conditions (cognitive reflection prompts, information sharing) to test mitigation.","core_discovery":"The paper's central claim is that LLM agents are not neutral optimizers in inventory tasks: under uncertainty, they systematically deviate from the optimal order quantity in the same direction as humans, and the deviations scale with the structure of the supply chain. AIM-Bench operationalizes this by running agents through a series of replenishment decisions in single-stage and multi-stage chains, measuring pull-to-center and bullwhip against normative baselines. The key result is that these biases are model-dependent but consistently present, and that cognitive reflection and information sharing reduce them, suggesting the biases respond to deliberate intervention rather than being fixed.","pith_inferences":["Because the benchmark mirrors classic newsvendor and beer-game setups, the same biases may appear in other operational tasks like pricing or capacity planning where humans also anchor and overreact.","The mitigation results suggest a testable field experiment: in a live supply chain, adding a reflection step or a demand-sharing dashboard should shrink order variance; if it does not, the benchmark effect may not transfer beyond text-based tasks.","The benchmark could be extended to multi-echelon networks with more than two stages or with perishable goods, where bias costs are higher, to see whether cognitive reflection still helps.","A falsifiable corollary is that pull-to-center magnitude will correlate with the model's calibration on demand distributions; better-calibrated models should show less bias."],"forward_implications":["LLM-based inventory agents will exhibit order distortions similar to human planners in the same task, so real deployments must budget for suboptimal ordering under uncertainty.","The benchmark provides quantitative bias measures (pull-to-center and bullwhip magnitudes) that can be compared across models, enabling model selection for operations roles.","Prompting agents to reflect before ordering reduces bias, implying prompt design is a cheap lever for safer AI operations.","Sharing demand information across chain stages reduces bullwhip, implying information architecture matters as much as model choice.","Different LLMs show different degrees of bias, so screening for bias should become a step in deploying agents for inventory management."],"supporting_citations":[],"fun_headline_variants":["LLM inventory agents order like biased humans","AI inventory managers show human-like ordering bias","AIM-Bench reveals LLM agents' pull-to-center and bullwhip","Cognitive reflection shrinks LLM inventory bias, AIM-Bench finds"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The measured biases are intrinsic to the LLMs' decision policies rather than artifacts of the benchmark's framing — for instance, how the demand numbers are presented or how order options are phrased.","fun_headline_variants_meta":{"raw":{"variants":["LLM inventory agents order like biased humans","AI inventory managers show human-like ordering bias","AIM-Bench reveals LLM agents' pull-to-center and bullwhip","Cognitive reflection shrinks LLM inventory bias, AIM-Bench finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000627,"raw_usage":{"total_tokens":2742,"prompt_tokens":751,"completion_tokens":1991,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":1923}},"tokens_in":495,"tokens_out":1991,"duration_ms":18285,"temperature":1.0,"reasoning_tokens":1923,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:55:41.534363+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same replenishment experiments with the same expected costs but with demand expressed in different units or scales; if the pull-to-center effect shifts substantially with the frame, the claimed human-like bias is partly a prompt artifact, not a stable property of the models.","supporting_citations":[],"review_version":1}