Pith. sign in

REVIEW 4 major objections 4 minor 63 references

Multi-Agent Reinforcement Learning for Dynamic Pricing in Supply Chains: Benchmarking Strategic Agent Behaviours under Realistically Simulated Market Conditions

T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read In a market simulation built from real retail transactions, multi-agent reinforcement learning price-setters out-earn rule-based pricing by a wide margin — the MADQN variant most of all — while eroding fairness and price stability.

desk verdict A legitimate MARL pricing benchmark with an honest limitations section, but the headline revenue result is a simulator artifact driven by a near-zero-elasticity demand curve, not a property of MARL. read the letter →

arxiv 2507.02698 v1 pith:Q3L5CUG3 submitted 2025-07-03 cs.LG econ.EM

classification cs.LGecon.EM
keywords multi-agentreinforcementlearningdynamicpricingsupplychainsimulationdemandforecastingLightGBMMADQNMADDPGQMIX
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that multi-agent reinforcement learning (MARL) can beat the static, rule-based pricing logic built into enterprise systems, and that the gains come with a measurable cost. In a two-year weekly simulation of a retail market, built on real transaction data and a gradient-boosted demand model, trained pricing agents far out-earned rule-based competitors: the MADQN variant averaged about forty-four times the rule-based return, with QMIX and MADDPG also far ahead. The paper's point is not just that learning beats rules, but that it does so through an emergent strategy — exploiting the simulated demand's near-zero sensitivity to price — which shows up in the numbers as lower fairness, higher price volatility, and almost no price convergence. Rule-based agents were nearly perfectly fair and stable but never competed; hybrids that mixed learning and rule-based agents kept most of the revenue while restoring stability. If the paper is right, the practical lesson is that adaptive pricing's value depends on market composition and on how price-sensitive demand really is.

What carries the argument

The argument runs on four pieces of machinery. First, the demand model: a LightGBM gradient-boosted regressor trained on weekly, log-transformed demand from the Online Retail II dataset (test $R^2 = 0.74$), which turns each agent's chosen price into a sales quantity and thereby defines the physics of the market; a counterfactual price-scaling analysis on this model yields the near-zero elasticity ($\varepsilon = -0.072$) that later lets agents profit from price increases. Second, the simulation environment: a weekly-stepped market running 104-week episodes for 30 episodes, in which four agents with identical five-product portfolios compete, with rewards combining revenue change and a quadratic price-instability penalty. Third, the MARL algorithms: MADQN (independent deep Q-learning over discrete price changes of $-10\%$ to $+10\%$), MADDPG (continuous actor-critic with centralized critics and decentralized execution), and QMIX (per-agent Q-networks combined by a monotonic mixing network that enforces consistency between individual and joint action values). Fourth, the rule-based baselines — static markup, competitor matching, historical anchor, demand responsive, and seasonal pricing — plus the structural metrics (Jain's fairness index, price volatility, Nash-equilibrium proximity, price convergence, optimality gap) that expose the trade-off the revenue numbers hide.

What would settle it

A decisive check is to re-run the identical agent zoo inside a market whose demand law is known and elastic — say, a simulator with true price elasticity $\varepsilon = -1.5$ and cross-elasticity that redistributes demand toward the cheapest competitor — and see whether MADQN's revenue edge over rule-based pricing survives or collapses. A cheaper, data-only check: take product-weeks from the Online Retail II data where genuine price changes occurred, and test whether the demand model's out-of-distribution predictions at counterfactual prices match the sales actually observed; if true demand responds far more than $\varepsilon = -0.072$, the revenue ordering is an artifact of the fitted model rather than a property of MARL.

Watch

Extended reading notes

Core claim

The paper's central claim is that MARL pricing agents — MADQN in particular — far outperform rule-based pricing agents in revenue within its simulated market, and that this outperformance is emergent strategic behaviour rather than a fixed property of any single algorithm. Across eight independent runs, an all-MADQN market produced mean per-agent revenue of £997,669 against £22,817 for an all-rule-based market (a 4,272.5% increase), with QMIX at £393,121 and MADDPG at £89,861. The same experiments quantify the trade-off: MADQN scored the lowest fairness on Jain's index (0.5844 versus 0.9896 for rule-based), the highest price volatility (0.085 versus 0.024), the highest market-share volatility (22.4 percentage points), and almost no price convergence. The paper attributes this pattern to agents learning to exploit the demand model's inelasticity (estimated price elasticity $\varepsilon = -0.072$): they raise prices because demand barely falls, which is optimal in the simulation but would not transfer to elastic markets.

Load-bearing premise

The revenue ranking rests entirely on the assumption that the LightGBM demand model correctly predicts what customers would buy at prices far outside its training data — the model says demand barely moves with price, so the agents' gains largely come from 'raise the price, sales barely drop,' and if real demand reacted more strongly, the ranking could reverse.

Editorial extensions

If this is right

  • If the central claim holds, adopters of MARL-based dynamic pricing should expect revenue gains to be concentrated in low-elasticity product lines, because the agents learn to exploit price-insensitive demand rather than to match customer willingness to pay.
  • Market composition is a design lever: hybrid populations that mix learning agents with rule-based agents preserve most of the revenue advantage while recovering near-rule-based fairness and price convergence, so deployment decisions are about the mix, not just the algorithm.
  • The fairness and volatility metrics used here give operators concrete monitoring signals: a market dominated by aggressive learners shows low Jain's index, high price volatility, and near-zero convergence, which read as early warnings rather than acceptable side effects.
  • The advantage shrinks where demand is elastic, where competitors respond strongly to price, or where regulation constrains price-setting — exactly the settings the paper lists as future work.
  • The authors conclude that MARL is most beneficial when demand is predictable, competition is strategic, and price setting is flexible, which makes the simulation a best-case scenario for learning agents rather than a general proof.
  • The same mechanism that produces the revenue — coordinated-looking price increases against inelastic demand — is the pattern antitrust scrutiny would flag as algorithmic coordination; the paper documents the fairness cost but leaves the regulatory implication unstated.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own significance test did not reach the 5% threshold (Wilcoxon $p = 0.125$ across four-agent runs), so the revenue ordering is best read as a directional result; a natural extension is re-running the benchmark with 8–16 agents per market, which would also enrich the competitive dynamics being measured.
  • Because the demand model is nearly price-invariant, the simulation leans toward a trivial optimum: raise prices to the top of the action space. Re-running the same agents against a demand law with realistic elasticity (say $\varepsilon \approx -1$ to $-2$, with cross-elasticity that shifts demand to cheaper competitors) would test whether MADQN's dominance survives when price actually moves demand
  • An implicit operational reading the paper does not develop: the same metrics that expose the trade-off (Jain's index, volatility, convergence) could serve as live guardrails in a deployed system, switching a market toward hybrid or rule-based pricing when fairness or stability thresholds are breached.
  • The mechanism behind MADQN's revenue — coordinated-looking price increases against inelastic demand — is also the pattern antitrust scrutiny would flag as algorithmic coordination; the paper documents the fairness cost but leaves the regulatory implication unstated.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper presents a multi-agent reinforcement learning (MARL) benchmark for dynamic pricing in a supply-chain setting. The authors use the UCI Online Retail II dataset, train a LightGBM demand model, and build a weekly simulation environment in which MADDPG, MADQN, and QMIX agents compete against rule-based pricing baselines. They report that MARL agents, particularly MADQN, achieve substantially higher revenue than rule-based agents, but with lower fairness and higher price volatility, and they interpret this as evidence of emergent strategic behavior. The paper also introduces hybrid agent configurations and evaluates coordination and welfare metrics.

Significance. If the revenue ranking were credible, the paper would provide a useful benchmark for MARL-based dynamic pricing and a concrete illustration of the trade-off between revenue, fairness, and stability. The authors deserve credit for building a reproducible simulation pipeline with a real-world transaction dataset, a pre-trained demand model, multiple MARL algorithms, and a broad set of market-level metrics. However, the central quantitative claim is undermined by two load-bearing problems: the reported statistical test does not support the word 'significantly,' and the demand model's near-zero price elasticity, combined with the authors' own admission of loophole-driven price exploitation, makes the revenue comparisons unreliable as evidence of strategic behavior. The contribution is therefore best viewed as an environment and benchmark design, not as a validated finding about MARL advantages.

major comments (4)
  1. [Section 4, Table 4 and Section 6] The central claim that 'MADQN in particular, significantly outperformed rule-based agents in terms of revenue' (Section 6) is directly contradicted by the paper's own statistical result: the Wilcoxon signed-rank test comparing MARL-only configurations to the rule-based baseline yields p = 0.125 with only four agents. The text in Section 4 acknowledges this lack of significance, but the conclusion restates the claim without the caveat. This is a load-bearing inconsistency that must be fixed, either by collecting more runs/agents to achieve adequate power or by explicitly reframing the result as a numerical improvement that is not statistically significant.
  2. [Section 3.3.2 and Section 5.3.1] The demand model's price elasticity is estimated by scaling prices 0.5x to 2.5x on the test set while holding other features fixed, but the model was trained on historical transaction prices only. No validation is provided that the model's predictions remain accurate for the price levels actually selected by the MARL agents, which lie outside the training distribution. Because Section 5.3.1 concedes that MADQN and QMIX 'exploited this inelasticity by raising prices,' the very large revenue advantages in Table 4 may be an artifact of extrapolating a near-zero elasticity curve rather than a property of the agents' strategic behavior. Please provide out-of-sample or counterfactual validation for the price range used by the agents, or restrict the action space to prices that are within the support of the training data.
  3. [Section 5.3.1 and Abstract] The authors describe the high-price strategies as 'loophole-driven,' yet the abstract and conclusion present the revenue gains as evidence of 'emergent strategic behaviour not captured by static pricing rules.' If the dominant strategy is to raise prices into an extrapolated region of the demand curve, the observed behavior is better characterized as exploitation of a model artifact than as strategic adaptation. The manuscript needs to disentangle these two interpretations and, if the loophole interpretation is correct, substantially weaken the claims of emergent strategic behavior.
  4. [Section 4, Table 5] Several headline market-level metrics—Nash Equilibrium Proximity (NEP), Price Convergence (PC), Revenue Optimality Gap (ROG), and Welfare Fairness (WF)—are reported as point estimates without variance or confidence intervals, while Jain's Fairness and Market Volatility are reported with ± ranges. Without error bars, the paper cannot support statements such as '4x MADQN shows the lowest Nash equilibrium proximity' or the comparisons of coordination across configurations. These metrics should be reported with the same run-level variability as the other columns, presumably over the eight independent simulation runs mentioned in Section 4.
minor comments (4)
  1. [Section 4] The word 'reutrns' in the sentence 'To assess performance differences in per-agent reutrns' is a typo and should read 'returns.'
  2. [Table 5] Some table entries contain stray spaces inside numeric values, for example '0 .9451', '22 .4', and '0 .5788'; these should be cleaned up for readability.
  3. [Table 3] The column header 'Stabil.' is abbreviated without a definition; the table would be clearer if the full metric name 'Stability' were used or a footnote were added.
  4. [Section 5.3.3] The sentence 'This feature heavily is shaped by local economic conditions and retail habits of that period' contains a word-order error; it should read 'This feature is heavily shaped by local economic conditions...'.

Circularity Check

1 steps flagged · score 6.0 of 10

MADQN's revenue supremacy is a mechanical consequence of the fitted near-zero-elasticity demand curve plus the revenue reward, not an independent MARL result.

  1. fitted input called prediction [Section 3.3.2 (price sensitivity), Section 4/Table 4 (revenue ranking), Section 5.3.1 (loophole admission), Appendix D (Revenue per Agent)]
    "The resulting ε =−0.072 indicates inelastic demand, which is typical for giftware products [56]... Several MARL agents, including MADQN and QMIX, exploited this inelasticity by raising prices across episodes, knowing demand would remain fairly stable. While optimal in this context, such behaviour would likely be restricted in markets where elasticity burdens pricing power."

    The demand model is fitted on historical transaction prices, and the elasticity used to interpret the simulation is measured by scaling prices through that same model while holding other inputs fixed. The simulator then scores agent-chosen prices with the same fitted model, and revenue is defined as R_agent = Σ P_t·Q_t (Appendix D). With ε = -0.072, predicted quantity is nearly constant in price, so revenue is nearly proportional to price; any revenue-rewarded agent is driven to the upper price bound. The central claim that MADQN 'significantly outperformed rule-based agents in terms of revenue' is therefore a direct consequence of the fitted demand curve plus the reward function, not an emergent property of MARL.

full rationale

No self-citation chain or definitional equation loop is present: the LightGBM model is trained before the simulations, and the MARL algorithms are standard implementations. The main circularity is partial and located in the revenue-ranking claim. Because the fitted demand model has near-zero price elasticity and the agents optimize revenue, the simulation mechanically rewards price increases; the paper admits MADQN and QMIX exploited this inelasticity. Thus the headline revenue advantage is substantially determined by the fitted input and the reward, not by an external benchmark or by strategic interactions that are robust to demand elasticity. The fairness, price-stability, and hybrid-configuration findings have independent content and are not circular. Also, the paper concedes the revenue comparison was not statistically significant (Wilcoxon p = 0.125), and its conclusion uses 'significantly' loosely; that is a correctness/statistical issue, not circularity. Overall, the central revenue result is partially circular, while the rest of the study is self-contained; score 6 reflects that the prediction reduces by construction to the fitted elasticity and reward design.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central results rest on the fitted LightGBM demand model, the hand-selected simulation configuration, and the reward design. The demand model's fitted elasticity is the main driver of the revenue ranking, and its counterfactual validity is assumed rather than demonstrated. No new physical or theoretical entities are introduced; the 'agents' are standard RL algorithm instances in a custom simulation.

free parameters (5)
  • Demand model price elasticity = epsilon = -0.072
    Derived from LightGBM counterfactual price scaling on the test set; it is the main driver of revenue outcomes and is not externally validated.
  • LightGBM hyperparameters = n_estimators=2048, learning_rate=0.03, num_leaves=256, early stopping patience=100
    Chosen by tuning; the fitted model is the environment's ground truth for demand in all simulations.
  • Semantic cluster count K = 20
    Selected by hand to balance category granularity; affects product state features and agent portfolios.
  • MARL reward price-stability penalty weights = not reported
    Appendix diagrams say rewards combine revenue change with a quadratic price-instability penalty, but coefficient values are not given; this directly shapes the reported stability and volatility scores.
  • MADQN/QMIX discrete action range = -10% to +10%
    Discrete price adjustment space chosen by hand for the value-based agents; restricts the set of achievable pricing strategies.
assumptions (4)
  • domain assumption The LightGBM model trained on historical prices gives valid counterfactual demand predictions for agent-chosen prices far outside the training distribution (0.5x to 2.5x price scaling).
    Section 3.3.2 uses counterfactual scaling to compute elasticity, and the simulation uses these predictions as ground truth for agent revenue.
  • ad hoc to paper The fitted demand model's near-zero elasticity is representative of the studied market and not an artifact of extrapolation.
    Section 3.3.2 reports epsilon = -0.072 and Section 5.3.1 acknowledges agents exploit this inelasticity; no external validation of the counterfactual demand response is provided.
  • domain assumption A weekly, four-agent, five-product simulation with no inventory or supply constraints captures the relevant supply-chain pricing dynamics.
    Section 3.7.1 sets 104-week episodes, four agents, and five products; Section 5.3.2 and 5.3.4 note the absence of supply limits and inventory-sensitive pricing.
  • ad hoc to paper The reward signal, revenue plus a quadratic price-stability penalty, reflects the business objective for pricing agents.
    Appendix E describes the reward; the penalty weight is not specified, and it directly determines the reported stability and volatility trade-offs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multi-Agent Reinforcement Learning for Dynamic Pricing in Supply Chains: Benchmarking Strategic Agent Behaviours under Realistically Simulated Market Conditions." pith.science (2026). https://pith.science/paper/Q3L5CUG3

@misc{pith2026250702698,
  author       = {Pith},
  title        = {Pith review of: Multi-Agent Reinforcement Learning for Dynamic Pricing in Supply Chains: Benchmarking Strategic Agent Behaviours under Realistically Simulated Market Conditions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q3L5CUG3}},
  note         = {Machine review of arXiv:2507.02698}
}
read the original abstract

This study investigates how Multi-Agent Reinforcement Learning (MARL) can improve dynamic pricing strategies in supply chains, particularly in contexts where traditional ERP systems rely on static, rule-based approaches that overlook strategic interactions among market actors. While recent research has applied reinforcement learning to pricing, most implementations remain single-agent and fail to model the interdependent nature of real-world supply chains. This study addresses that gap by evaluating the performance of three MARL algorithms: MADDPG, MADQN, and QMIX against static rule-based baselines, within a simulated environment informed by real e-commerce transaction data and a LightGBM demand prediction model. Results show that rule-based agents achieve near-perfect fairness (Jain's Index: 0.9896) and the highest price stability (volatility: 0.024), but they fully lack competitive dynamics. Among MARL agents, MADQN exhibits the most aggressive pricing behaviour, with the highest volatility and the lowest fairness (0.5844). MADDPG provides a more balanced approach, supporting market competition (share volatility: 9.5 pp) while maintaining relatively high fairness (0.8819) and stable pricing. These findings suggest that MARL introduces emergent strategic behaviour not captured by static pricing rules and may inform future developments in dynamic pricing.

Figures

Figures reproduced from arXiv: 2507.02698 by the authors.

Figure 1
Figure 1. Smoothed price–demand curve Price elasticity of demand (𝜀) was calculated as: 𝜀 = Δ𝑄/𝑄 Δ𝑃/𝑃 (1) The resulting 𝜀 = −0.072 indicates inelastic demand, which is typ￾ical for giftware products [56]. Although trained on log-transformed demand, elasticity was computed in the original scale using inverse￾transformed demand predictions. 3.4 Simulation Environment Design A custom simulation environment was developed to simul… view at source ↗
Figure 2
Figure 2. Final Episode Market Share per Agent across Con [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Distribution of price volatility measured by coefficient of variation (CV). [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: MADDPG Per-Agent Architecture. Each agent selects continuous price changes using a local Actor Network informed [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: MADQN Per-Agent Architecture. Each agent independently learns discrete pricing actions through a local Q-Network [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: QMIX Per-Agent Architecture. Each pricing agent uses a local Q-Network to learn discrete price adjustments based on [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

63 extracted references · 35 canonical work pages

  1. [1]

    Snellius: de Nationale Supercomputer

    2025. Snellius: de Nationale Supercomputer. https://www.surf.nl/snellius-de- nationale-supercomputer. Accessed: 2025-05-16

  2. [2]

    Khaled Abdalgader, Atheer A Matroud, and Khaled Hossin. 2024. Experimen- tal study on short-text clustering using transformer-based semantic similarity measure. PeerJ Computer Science 10 (2024), e2078. doi:10.7717/peerj-cs.2078

  3. [3]

    Anthony B Atkinson. 1999. The contributions of Amartya Sen to welfare economics. The Scandinavian Journal of Economics 101, 2 (1999), 173–190. doi:10.1111/1467-9442.00151

  4. [4]

    Kasun Bandara, Christoph Bergmeir, and Slawek Smyl. 2020. Forecasting across time series databases using recurrent neural networks on groups of similar series: A clustering approach. Expert Systems with Applications 140 (2020), 112896. doi:10.1016/j.eswa.2019.112896

  5. [5]

    Souhaib Ben Taieb, Gianluca Bontempi, Amir F Atiya, and Antti Sorjamaa. 2012. A review and comparison of strategies for multi-step ahead time series forecasting based on the NN5 forecasting competition. Expert Systems with Applications 39, 8 (2012), 7067–7083. doi:10.1016/j.eswa.2012.01.039 9

  6. [6]

    Maurício F Blos and Paulo E Miyagi. 2015. Modeling the supply chain disruptions: A study based on the supply chain interdependencies. IFAC-PapersOnLine 48, 3 (2015), 2053–2058. doi:10.1016/j.ifacol.2015.06.391

  7. [7]

    Lucian Busoniu, Robert Babuska, and Bart De Schutter. 2008. A comprehensive survey of multiagent reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) 38, 2 (2008), 156–172. doi:10.1109/TSMCC.2007.913919

  8. [8]

    Gérard P Cachon and Pnina Feldman. 2010. Dynamic versus static pricing in the presence of strategic consumers. The Wharton School, University of Pennsylvania (2010)

Show all 63 references
  1. [9]

    Lidia Ceriani and Paolo Verme. 2012. The origins of the Gini index: extracts from Variabilità e Mutabilità (1912) by Corrado Gini. The Journal of Economic Inequality 10 (2012), 421–443. doi:10.1007/s10888-011-9188-x

  2. [10]

    Daqing Chen. 2012. Online Retail II [Dataset]. https://archive.ics.uci.edu/dataset/ 502/online+retail+ii. doi:10.24432/C5CG6D

  3. [11]

    Daqing Chen, Kun Guo, and George Ubakanma. 2015. Predicting customer profitability over time based on RFM time series.International Journal of Business Forecasting and Marketing Intelligence 2, 1 (2015), 1–18. doi:10.1504/IJBFMI.2015. 075325

  4. [12]

    Daqing Chen, Sai Laing Sain, and Kun Guo. 2012. Data mining for the online retail industry: A case study of RFM model-based customer segmentation using data mining. Journal of Database Marketing & Customer Strategy Management 19, 3 (2012), 197–208. doi:10.1057/dbm.2012.17

  5. [13]

    Le Chen, Alan Mislove, and Christo Wilson. 2016. An Empirical Analysis of Algo- rithmic Pricing on Amazon Marketplace. In Proceedings of the 25th International Conference on World Wide Web. ACM, 1339–1349. doi:10.1145/2872427.2883089

  6. [14]

    Yixian Chen, Prakhar Mehrotra, Nitin Kishore Sai Samala, Kamilia Ahmadi, Viresh Jivane, Linsey Pang, Monika Shrivastav, Nate Lyman, and Scott Pleiman

  7. [15]

    Pritom Das, Tamanna Pervin, Biswanath Bhattacharjee, Md Razaul Karim, Nasrin Sultana, Md Sayham Khan, Md Afjal Hosien, and FNU Kamruzzaman. 2024. Optimizing real-time dynamic pricing strategies in retail and e-commerce using machine learning models. The American Journal of Eng...

  8. [16]

    Junfeng Dong, Beilei Rao, Yu Liu, Li Jiang, Wenxing Lu, and Qiang Guo. 2019. Pricing strategies for different periods during subsequent selling season for seasonal products. IEEE Access 8 (2019), 39479–39490. doi:10.1109/ACCESS.2019. 2953284

  9. [17]

    Raouya El Youbi, Fayçal Messaoudi, and Manal Loukili. 2023. Machine learning- driven dynamic pricing strategies in E-commerce. In 2023 14th International Conference on Information and Communication Systems (ICICS) . IEEE, 1–5. doi:10. 1109/ICICS60529.2023.10330541

  10. [18]

    Lawrence Emma. 2024. Enterprise Resource Planning (ERP) Systems for Streamlining Organizational Processes. Unpublished Manuscript (2024). https://www.researchgate.net/publication/386382658_Enterprise_Resource_ Planning_ERP_Systems_for_Streamlining_Organizational_Processes

  11. [19]

    Foerster, Yannis M

    Jakob N. Foerster, Yannis M. Assael, Nando de Freitas, and Shimon White- son. 2016. Learning to Communicate with Deep Multi-Agent Reinforcement Learning. In Advances in Neural Information Processing Systems , Vol. 29. Cur- ran Associates, Inc., 2137–2145. https://proceedings.n...

  12. [20]

    Kallirroi Georgila, Claire Nelson, and David Traum. 2014. Single-agent vs. multi- agent techniques for concurrent reinforcement learning of negotiation dialogue policies. In Proceedings of the 52nd Annual Meeting of the Association for Com- putational Linguistics (Volume 1: Lo...

  13. [21]

    Rajan Gupta and Chaitanya Pathak. 2014. A machine learning framework for predicting purchase by online customers based on dynamic pricing. Procedia Computer Science 36 (2014), 599–605. doi:10.1016/j.procs.2014.09.060

  14. [22]

    Adnène Hajji, Robert Pellerin, Pierre-Majorique Léger, Ali Gharbi, and Gilbert Babin. 2012. Dynamic pricing models for ERP systems under network externality. International Journal of Production Economics 135, 2 (2012), 708–715. doi:10.1016/ j.ijpe.2011.10.004

  15. [23]

    Ye Han, Xuefei Zhang, Jian Zhang, Qimei Cui, Shuo Wang, and Zhu Han. 2019. Multi-agent reinforcement learning enabling dynamic pricing policy for charging station operators. In 2019 IEEE Global Communications Conference (GLOBECOM) . IEEE, 1–6. doi:10.1109/GLOBECOM38437.2019.9013999

  16. [24]

    Hendricks and Kate W

    Walter A. Hendricks and Kate W. Robey. 1936. The Sampling Distribution of the Coefficient of Variation. Annals of Mathematical Statistics 7, 3 (1936), 129–132. doi:10.1214/aoms/1177732503

  17. [25]

    Jacob Hilton, Jie Tang, and John Schulman. 2023. Scaling laws for single-agent reinforcement learning. arXiv preprint arXiv:2301.13442 (2023). doi:10.48550/ arXiv.2301.13442

  18. [26]

    Andreas Hinterhuber. 2004. Towards value-based pricing—An integrative frame- work for decision making. Industrial Marketing Management 33, 8 (2004), 765–778. doi:10.1016/j.indmarman.2003.10.006

  19. [27]

    Samuel B Hwang and Sungho Kim. 2006. Dynamic pricing algorithm for E- Commerce. In Advances in Systems, Computing Sciences and Software Engineering: Proceedings of SCSS05. Springer, 149–155. doi:10.1007/1-4020-5263-4_24

  20. [28]

    Rajendra K Jain, Dah-Ming W Chiu, and William R Hawe. 1984. A Quantitative Measure of Fairness and Discrimination . Technical Report TR-301. Digital Equip- ment Corporation, Hudson, MA. https://www.cs.wustl.edu/~jain/papers/ftp/ fairness.pdf

  21. [29]

    Vipul Jain and Lyes Benyoucef. 2008. Managing long supply chain networks: some emerging issues and challenges. Journal of Manufacturing Technology Management 19, 4 (2008), 469–496. doi:10.1108/17410380810869923

  22. [30]

    Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Advances in Neural Information Processing Systems , Vol. 30. Curran Associates, Inc., 3146–3154. https://pro...

  23. [31]

    Byung-Gook Kim, Yu Zhang, Mihaela Van Der Schaar, and Jang-Won Lee. 2015. Dynamic pricing and energy consumption scheduling with reinforcement learn- ing. IEEE Transactions on Smart Grid 7, 5 (2015), 2187–2198. doi:10.1109/TSG. 2015.2495145

  24. [32]

    Konda and John N

    Vijay R. Konda and John N. Tsitsiklis. 2003. On actor-critic algorithms. SIAM Journal on Control and Optimization 42, 4 (2003), 1143–1166. doi:10.1137/ S0363012901385691

  25. [33]

    Oh Byung Kwon and J.J. Lee. 2001. A multi-agent intelligent system for efficient ERP maintenance. Expert Systems with Applications 21, 4 (2001), 191–202. doi:10. 1016/S0957-4174(01)00039-2

  26. [34]

    Derek Li, Andrew Jacobsen, and Adam White. 2021. Revisiting Experience Replay in Non-Stationary Environments. In Proceedings of the Adaptive and Learning Agents Workshop (ALA). https://ala2021.vub.ac.be/papers/ALA2021_paper_51. pdf

  27. [35]

    Le Li, Xiao Lin, Rudy R Negenborn, and Bart De Schutter. 2015. Pricing intermodal freight transport services: A cost-plus-pricing strategy. InComputational Logistics: 6th International Conference, ICCL 2015, Delft, The Netherlands, September 23-25, 2015, Proceedings 6. Springe...

  28. [36]

    Jiaxin Liang, Haotian Miao, Kai Li, Jianheng Tan, Xi Wang, Rui Luo, and Yueqiu Jiang. 2025. A Review of Multi-Agent Reinforcement Learning Algorithms. Electronics 14, 4 (2025), 820. doi:10.3390/electronics14040820

  29. [37]

    Jiaxi Liu, Yidong Zhang, Xiaoqing Wang, Yuming Deng, and Xingyu Wu. 2019. Dynamic Pricing on E-commerce Platform with Deep Reinforcement Learning: A Field Experiment. arXiv preprint arXiv:1912.02572 (2019). doi:10.48550/arXiv. 1912.02572

  30. [39]

    Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch

  31. [40]

    Arun K. Menon. 2024. Integrating Pricing Optimization Models with ERP Systems to Enhance Profitability in U.S. E-commerce Supply Chains. Global Journal of Engineering and Technology Advances 20, 2 (2024), 242–255. doi:10.30574/gjeta. 2024.20.2.0151 Open access under CC BY 4.0

  32. [41]

    Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Sha...

  33. [42]

    https://github.com/openai/maddpg Archived codebase provided as-is, based on the NeurIPS 2017 paper

    Multi-Agent Deep Deterministic Policy Gradient (MADDPG) - GitHub repository. https://github.com/openai/maddpg Archived codebase provided as-is, based on the NeurIPS 2017 paper

  34. [43]

    Gonçalo Neto. 2005. From Single-Agent to Multi-Agent Reinforcement Learning: Foundational Concepts and Methods. https://users.cs.utah.edu/~tch/CS6380/ resources/Neto-2005-RL-MAS-Tutorial.pdf Learning Theory Course 2005; 2

  35. [44]

    Vincent R Nijs, Shuba Srinivasan, and Koen Pauwels. 2007. Retail-price drivers and retailer profits. Marketing Science 26, 4 (2007), 473–487. doi:10.1287/mksc. 1060.0205

  36. [45]

    John F Nash Jr. 1950. Equilibrium Points in N-Person Games. Proceedings of the National Academy of Sciences 36, 1 (1950), 48–49. doi:10.1073/pnas.36.1.48

  37. [46]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCN...

  38. [47]

    Lei Ren, Xiaoyang Fan, Jin Cui, Zhen Shen, Yisheng Lv, and Gang Xiong. 2022. A multi-agent reinforcement learning method with route recorders for vehicle rout- ing in supply chain management. IEEE Transactions on Intelligent Transportation Systems 23, 9 (2022), 16410–16420. do...

  39. [48]

    Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Far- quhar, Jakob Foerster, and Shimon Whiteson. 2018. QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning. InProceed- ings of the 35th International Conference on Machi...

  40. [49]

    Shahriar Shafiee and Erkan Topal. 2010. An overview of global gold market and gold price forecasting. Resources Policy 35, 3 (2010), 178–189. doi:10.1016/j. resourpol.2010.05.004

  41. [50]

    Ravid Shwartz-Ziv and Amitai Armon. 2022. Tabular Data: Deep Learning is Not All You Need. Information Fusion 81 (2022), 84–90. doi:10.1016/j.inffus.2021.11.011

  42. [51]

    Melrose Roderick, James MacGlashan, and Stefanie Tellex. 2017. Implementing the deep q-network. arXiv preprint arXiv:1711.07478 (2017). doi:10.48550/arXiv. 1711.07478

  43. [52]

    Hussein Kamaldeen Smith. 2024. Dynamic Pricing and E-supply Chain Coor- dination: Effects on Inventory Optimization and Profit Margins. Unpublished Manuscript (2024). https://www.researchgate.net/publication/384805839_ Dynamic_Pricing_and_E-supply_Chain_Coordination_Effects_on...

  44. [53]

    Bo Sun, Hossein Nekouyan Jazi, Xiaoqi Tan, and Raouf Boutaba. 2024. Static Pric- ing for Online Selection Problem and its Variants.arXiv preprint arXiv:2410.07378 (2024). doi:10.48550/arXiv.2410.07378

  45. [54]

    Graves, Douglas A

    Rina Singh, Jeffrey A. Graves, Douglas A. Talbert, and William Eberle. 2018. Prefix and suffix sequential pattern mining. In Advances in Data Mining. Applications and Theoretical Aspects. Springer, 309–324. doi:10.1007/978-3-319-95786-9_24

  46. [55]

    Qiaochu Wang, Yan Huang, Param Vir Singh, and Kannan Srinivasan. 2023. Algorithms, Artificial Intelligence and Simple Rule Based Pricing.SSRN Electronic Journal (2023). doi:10.2139/ssrn.4144905

  47. [56]

    Sherry Shi Wang and Ralf Van Der Lans. 2018. Modeling gift choice: The effect of uncertainty on price sensitivity. Journal of Marketing Research 55, 4 (2018), 524–540. doi:10.1509/jmr.16.0453

  48. [57]

    J Michael Tarn, David C Yen, and Marcus Beaumont. 2002. Exploring the ratio- nales for ERP and SCM integration. Industrial Management & Data Systems 102, 1 (2002), 26–34. doi:10.1108/02635570210414631

  49. [58]

    Rafał Weron. 2014. Electricity price forecasting: A review of the state-of-the-art with a look into the future. International Journal of Forecasting 30, 4 (2014), 1030–1081. doi:10.1016/j.ijforecast.2014.08.008

  50. [59]

    Annie Wong, Thomas Bäck, Anna V Kononova, and Aske Plaat. 2023. Deep mul- tiagent reinforcement learning: Challenges and directions. Artificial Intelligence Review 56, 6 (2023), 5023–5056. doi:10.1007/s10462-022-10299-x

  51. [60]

    Christopher JCH Watkins and Peter Dayan. 1992. Q-learning. Machine Learning 8, 3 (1992), 279–292. doi:10.1007/BF00992698

  52. [61]

    Ziyuan Zhou, Guanjun Liu, and Ying Tang. 2023. Multi-agent reinforcement learning: Methods, applications, visionary prospects, and challenges. arXiv preprint arXiv:2305.10091 (2023). doi:10.48550/arXiv.2305.10091 11 Appendices A DATASET DESCRIPTION Table 6 provides an overview...

  53. [63]

    2016.Automated Pricing Agents in the On-Demand Economy

    Tony Wu, Anthony D Joseph, and Stuart J Russell. 2016.Automated Pricing Agents in the On-Demand Economy . Master’s thesis. University of California, Berkeley. https://www2.eecs.berkeley.edu/Pubs/TechRpts/2016/EECS-2016-57.html

  54. [2017]

    In Advances in Neural Information Processing Systems , Vol

    Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environ- ments. In Advances in Neural Information Processing Systems , Vol. 30. Curran Associates, Inc., 6379–6390. https://proceedings.neurips.cc/paper/2017/file/ 68a9750337a418a86fe06c1991a1d64c-Paper.pdf

  55. [2021]

    INFORMS Journal on Applied Analytics 51, 1 (2021), 76–89

    A Multiobjective Optimization for Clearance in Walmart Brick-and-Mortar Stores. INFORMS Journal on Applied Analytics 51, 1 (2021), 76–89. doi:10.1287/ inte.2020.1065

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.