Pith. sign in

REVIEW 1 cited by

Dual-Agent Deep Reinforcement Learning for Dynamic Pricing and Replenishment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.21109 v1 pith:P26JTZVL submitted 2024-10-28 cs.LG econ.GNq-fin.EC

Dual-Agent Deep Reinforcement Learning for Dynamic Pricing and Replenishment

classification cs.LG econ.GNq-fin.EC
keywords pricingdecisiondemandlearningreplenishmentapproachdeepdifferent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We study the dynamic pricing and replenishment problems under inconsistent decision frequencies. Different from the traditional demand assumption, the discreteness of demand and the parameter within the Poisson distribution as a function of price introduce complexity into analyzing the problem property. We demonstrate the concavity of the single-period profit function with respect to product price and inventory within their respective domains. The demand model is enhanced by integrating a decision tree-based machine learning approach, trained on comprehensive market data. Employing a two-timescale stochastic approximation scheme, we address the discrepancies in decision frequencies between pricing and replenishment, ensuring convergence to local optimum. We further refine our methodology by incorporating deep reinforcement learning (DRL) techniques and propose a fast-slow dual-agent DRL algorithm. In this approach, two agents handle pricing and inventory and are updated on different scales. Numerical results from both single and multiple products scenarios validate the effectiveness of our methods.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Two-Timescale Hierarchical Reinforcement Learning for Resilient Operations

    stat.ML 2026-07 conditional novelty 6.0

    Synchronized two-timescale hierarchical PPO-style learning converges in average optimality gap at O(T^{-1/2}) (faster under market sharpness) and raises simulated used-car profits under joint shocks.