Pith. sign in

REVIEW 1 cited by

Online Pricing and Allocation with Demand Learning and Fulfillment Cost

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.18049 v3 pith:O5P74MJ2 submitted 2025-01-29 cs.LG math.OCstat.ML

Online Pricing and Allocation with Demand Learning and Fulfillment Cost

classification cs.LG math.OCstat.ML
keywords demandlearninginventorypriceallocationocsaaonlineadditive-accuracy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We study online learning for a seller that jointly chooses per-period inventory positions and a uniform price, then fulfills realized demand through a downstream allocation. The main difficulty is not only demand learning: the price shifts demand and reshapes the transportation LP, making the population objective globally non-convex and non-smooth. To solve this problem, we propose OCSAA, an algorithm that exploits demand observations through counterfactual translation and proposes joint (price, inventory) decisions through lower-confidence optimism. OCSAA admits a polynomial-time additive-accuracy implementation for rational-polytope inventory sets. We prove a high-probability $\widetilde O(\sqrt T)$ regret guarantee and establish a matching-in-$T$ information-theoretic lower bound. Our results illustrate an effective integration of statistical learning methodologies with complex operations research problems.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Two-Timescale Hierarchical Reinforcement Learning for Resilient Operations

    stat.ML 2026-07 conditional novelty 6.0

    Synchronized two-timescale hierarchical PPO-style learning converges in average optimality gap at O(T^{-1/2}) (faster under market sharpness) and raises simulated used-car profits under joint shocks.