Pith. sign in

REVIEW 2 cited by

Deep Inventory Management

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.03137 v3 pith:7N3P4VFZ submitted 2022-10-06 cs.LG math.OC

classification cs.LGmath.OC
keywords learninginventorycontrolreinforcementbackpropdeepdirectperiodic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work provides a Deep Reinforcement Learning approach to solving a periodic review inventory control system with stochastic vendor lead times, lost sales, correlated demand, and price matching. While this dynamic program has historically been considered intractable, our results show that several policy learning approaches are competitive with or outperform classical methods. In order to train these algorithms, we develop novel techniques to convert historical data into a simulator. On the theoretical side, we present learnability results on a subclass of inventory control problems, where we provide a provable reduction of the reinforcement learning problem to that of supervised learning. On the algorithmic side, we present a model-based reinforcement learning procedure (Direct Backprop) to solve the periodic review inventory control problem by constructing a differentiable simulator. Under a variety of metrics Direct Backprop outperforms model-free RL and newsvendor baselines, in both simulations and real-world deployments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hard Constraints, Smooth Gradients: Learning Feasible Inventory Policies via Differentiable Projection

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A differentiable QP-projection plus dual-aware rounding layer lets deep RL policies enforce interdependent hard constraints, achieving near-optimal cost on small instances and 2.5–3.2% savings on an ASML case study.

  2. Structure-Informed Deep Reinforcement Learning for Inventory Management

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A generic DirectBackprop deep RL policy, trained only on historical demand across many products, matches or beats classical inventory heuristics in five problem settings, and structural monotonicity penalties improve ...

Pith tools