Pith. sign in

REVIEW 4 minor 70 references

Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue

T0 review · 0 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper proves a pilot-corrected layered decision-partitioning policy attains the minimax-optimal regret rate for semiparametric contextual dynamic pricing with Hölder-smooth multimodal revenue, arbitrary covariate sequences, and bounded

desk verdict A clean, well-argued minimax result for smooth multimodal pricing; the known-smoothness assumption is the real limitation, not a flaw in the proof. read the letter →

arxiv 2608.03142 v1 pith:WOEQIZ4Z submitted 2026-08-04 stat.ML cs.AIcs.LG

classification stat.MLcs.AIcs.LG MSC 62G0562C20
keywords contextualdynamicpricingsemiparametricdemandboundedquantityfeedbackshape-freerevenueminimaxregretHöldersmoothnesslocalpolynomialregressionadaptiveexploration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper studies a seller who must price in sequence, seeing one context vector per buyer and then a bounded purchase quantity, which may be binary, discrete, or continuous. Demand is assumed to depend on price only through the surplus between a linear valuation index and the price, with both the index parameter and the demand-response function unknown; the induced demand link is only assumed Hölder-smooth with exponent $\beta$. The paper's central claim is that, even under arbitrary (possibly adversarial) context sequences and with revenue that may be multimodal with nonunique optimal prices, the minimax regret is $\widetilde{O}(T^{(\beta+1)/(2\beta+1)})$: the proposed policy achieves this rate up to log factors, and a matching lower bound shows no policy can do better in horizon dependence. The value of the claim is that smoothness alone, without strong unimodality, concavity, unique price optima, or distributional context assumptions, already determines the optimal horizon exponent.

What carries the argument

The load-bearing object is the pilot-corrected local-polynomial feature $\psi_{t,j}(w)$, which stacks the usual local-polynomial basis with products of derivative coefficients and the unknown valuation parameter $\theta_\star$, evaluated at a residual action $w$ in bin $I_j$. Writing the conditional mean as a linear function of these composite coefficients turns the otherwise nonconvex joint estimation of the index and the link into a convex ridge regression, leaves only an $O(h^\beta+\eta^2)$ approximation error, and makes the pilot-index error second-order. Permanent layer–bin labels keep the sampling predictable, so self-normalized concentration applies; a layered optimistic-elimination r

What would settle it

Go to the lower-bound family: take the smooth baseline link from the flat-revenue construction, place a perturbation of width $K_T^{-1}$ and height $K_T^{-\beta}$ in one of $K_T$ separated regions inside the flat interval, and compute the total KL divergence for $K_T \asymp T^{1/(2\beta+1)}$. If the KL is bounded away from $O(1)$, the testing argument behind the $\Omega(T^{(\beta+1)/(2\beta+1)})$ lower bound fails; if the proposed policy run on the unperturbed flat baseline shows regret $\omega(\log T)$, the layer-occupancy analysis fails.

Watch

Extended reading notes

Core claim

The discovery is a policy and a matching impossibility result. The policy is a pilot-corrected layered decision-partitioning (LDP) policy with an adaptive directional pilot that certifies the valuation index only along observed covariate directions; a local-polynomial regression whose augmented feature absorbs the first-order pilot index error into composite coefficients; permanent layer–bin labels assigned before demand is observed; and global optimistic elimination over the entire residual price domain. The paper proves this policy incurs $\widetilde{O}(T^{(\beta+1)/(2\beta+1)})$ expected regret for fixed problem primitives. It then constructs a constant-context, binary-demand subclass of

Load-bearing premise

If the induced demand link is not actually Hölder-smooth at the assumed order $\beta$ with the known constant $L_g$, the policy's local-polynomial bias and the whole regret rate $\widetilde{O}(T^{(\beta+1)/(2\beta+1)})$ are not guaranteed; the seller must treat $\beta$ and $L_g$ as known.

Editorial extensions

If this is right

  • At $\beta=1$, the general-rate statement recovers the $\widetilde{O}(T^{2/3})$ regret already known for Lipschitz links with arbitrary contexts, now for bounded nonbinary quantity feedback as well.
  • For twice-smooth demand ($\beta=2$), the rate is $\widetilde{O}(T^{3/5})$ without strong unimodality; existing smooth-context results at this exponent required strong unimodality, so the geometry assumption is not needed for this horizon exponent.
  • The lower bound shows the hard region is a flat-revenue interval with well-separated perturbations: the minimax difficulty is global mode discovery, not local optimization.
  • Because the bound holds for arbitrary covariate sequences, it applies when contexts are chosen adversarially or depend on past prices, up to log factors.
  • The same policy automatically handles binary purchase feedback as a special case, so the result is a strict broadening of the binary-feedback semiparametric pricing problem.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The prediction-lifting trick—absorbing $\theta_\star$ times derivative coefficients into composite coefficients—is not tied to pricing; it could plausibly be reused in other online single-index problems, such as contextual bandits with misspecified links, whenever only predictions at queried points are needed.
  • Since the lower bound already holds with a single constant context, stochastic contexts cannot improve the horizon exponent; any faster rate must come from extra assumptions such as strong unimodality or feature diversity, a point the paper leaves implicit but its construction implies.
  • A shape-adaptive policy that detects a unique quadratic revenue mode and switches to localization may beat this rate on strongly unimodal instances while retaining the guarantee over the unrestricted class; the paper explicitly leaves this best-of-both-worlds direction open.
  • Whether simultaneous adaptation to unknown smoothness order $\beta$ and unknown Hölder constant $L_g$ costs an extra factor is not resolved; the analysis treats both as known policy inputs, and the paper lists this as an open question.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The paper considers contextual dynamic pricing with a semiparametric surplus-index demand model: latent valuations are linear in covariates plus i.i.d. noise, and expected demand is an unknown function of the surplus p − x^T θ*. The model allows bounded, possibly nonbinary quantities, arbitrary context sequences, and no concavity/unimodality assumptions on revenue. The authors propose a pilot-corrected layered decision-partitioning (LDP) policy that (i) uses uncertainty-triggered uniform-price exploration to certify the valuation index only along the current covariate direction, (ii) absorbs the first-order pilot-index displacement into lifted local-polynomial coefficients, and (iii) performs global optimistic elimination over the residual price domain with permanent layer–bin labels. The main upper-bound result (Theorem 5.4) is Õ(T^{(β+1)/(2β+1)}) regret under Hölder smoothness β ≥ 1, and Theorem 5.6 gives a matching lower bound on a constant-context binary-demand subclass, establishing the minimax horizon exponent up to logarithmic factors for fixed smoothness parameters. The appendix contains the pilot confidence lemma, the lifted linearization with O(h^β + η^2) remainder, the uniform confidence event, layer-occupancy counting, the pilot-count determinant argument, and the lower-bound construction with a flat-revenue baseline and separated perturbations.

Significance. If the result holds, this is a substantial theoretical contribution. It extends the shape-free Lipschitz contextual pricing results (T 2/3 regret) to higher-order smoothness while removing the strong-unimodality structure used by recent smooth-pricing analyses, and it simultaneously handles arbitrary covariate sequences and bounded quantity feedback. The rate matches the nonparametric smooth-pricing benchmark of Wang et al. (2021), which is the natural target for the shape-free regime. The proof is unusually complete: the appendix provides the martingale/self-normalized concentration arguments, the deterministic misspecification bound for the pilot-corrected lifted regression, pathwise pilot-exploration control, and an explicit hard-instance family with a KL-based testing argument. I especially credit the permanent-label design, which cleanly preserves predictability under adaptive sampling, and the lower-bound construction, which genuinely places the difficulty in the nonparametric flat-optimum region rather than in the contextual parameter. The known-smoothness limitation is disclosed in the conclusion and is a scope restriction, not an internal inconsistency.

minor comments (4)
  1. [Assumption 3.4 and Section 6] Assumption 3.4 treats β and L_g as known, and the tuning in Theorem 5.4 (N and η) depends on β. The paper explicitly defers adaptation to unknown smoothness to future work in Section 6. This is a genuine scope limitation, not a load-bearing error, but the abstract's 'minimax-optimal' claim should be qualified as holding over the class with known smoothness parameters; a sentence to this effect in the introduction would prevent misreading.
  2. [Appendix A.1/A.2] The appendix labels 'Proof of Theorem 3.5', 'Proof of Theorem 4.1', and 'Proof of Theorem 4.2' refer, respectively, to Proposition 3.5, Lemma 4.1, and Proposition 4.2. These cross-reference mismatches should be corrected.
  3. [Corollary 5.7] The corollary states an expected-regret bound but does not spell out the conversion from the high-probability bound of Theorem 5.4. Since regret is deterministically bounded by BDT, the standard argument with δ ≍ 1/T applies; please include a brief sentence so the expectation bound is formally derived.
  4. [Lemma A.2, inequality (46)] The replacement of the confidence-radius terms by min(1, ‖ψ‖_{(Λ+λI)^{-1}}) factors implicitly uses that the constants Cψ and Cz are at least 1 and that ι_T > 1. These conditions are satisfied after enlarging constants as in the definitions (25) and (35), but the proof should state this explicitly; otherwise the displayed inequality appears to require an additional justification.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: upper and lower bounds are self-contained derivations from the stated model assumptions; self-citations are non-load-bearing.

full rationale

The paper's central derivation chain is not circular. The pilot-corrected local representation (Eq. 4) follows from a Holder Taylor expansion of the induced link g around bin anchors, with the lifted coefficient z_j defined so that the first-order pilot-index displacement is absorbed into the feature; the remainder bound O(h^beta + eta^2) is derived from the Holder condition and the pilot certificate, not assumed as the conclusion. The confidence radius (7) is obtained by combining self-normalized martingale concentration, ridge bias, and explicit approximation-error bounds, with constants constructed in the proof of Proposition 4.2. The layer occupancy and pilot-count bounds are pathwise elliptical-potential/determinant arguments. The regret rate in Theorem 5.4 is the result of balancing N, h, and eta, and no fitted quantity is renamed as a prediction. The lower bound (Theorem 5.6) explicitly constructs a flat-revenue baseline and statistically indistinguishable perturbations, verifies that these instances satisfy the model assumptions, and proves the matching horizon exponent via a KL/Bretagnolle-Huber argument; it does not import the upper bound. Assumption 3.4 (known beta and L_g) is a regularity premise on the object being learned and is explicitly acknowledged as a limitation with adaptation left to future work. The self-citations to Gong et al. (2025) are used only as an algorithmic antecedent for the LDP architecture, not as evidence for the new rate, and the present proofs are self-contained. No circular step was found.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on three domain assumptions (bounded zero-mean noise, surplus-only demand, known Holder smoothness) plus standard statistical inequalities. The algorithmic hyperparameters N, eta, and lambda are tuned to the horizon, not estimated from data, so they do not create a fitted-to-data circularity. No new physical or model entities are introduced.

free parameters (3)
  • Number of residual bins N = N = ceil(T^(1/(2beta+1)))
    Horizon-dependent tuning parameter balancing the O(sqrt(NT)) statistical cost across bins with the O(T h^beta) approximation cost. Not fitted to observations.
  • Pilot target accuracy eta = eta = sqrt(min{h^2, h^beta}) with h = 2B/N
    Chosen so the pilot-index error is second order and does not worsen the nonparametric rate. It is a horizon-dependent algorithmic choice, not a data-fitted parameter.
  • Ridge regularization lambda = any constant > 0 independent of T
    Fixed positive regularization used for convexity and self-normalized concentration. The analysis allows any such value; it is not fitted to the data.
assumptions (4)
  • domain assumption Zero-mean, i.i.d., bounded valuation shocks with support [-B_epsilon, B_epsilon], and x^T theta* in [B_epsilon, B - B_epsilon] for all x in X (Assumption 3.1).
    This guarantees valuations lie in [0,B], makes uniform-price pilot observations unbiased for the index, and provides the boundedness needed for martingale concentration.
  • domain assumption Surplus-dependent single-index demand: E[y_t | G_t or v_t] = q(v_t - p_t) for an unknown q, with q(z)=0 outside [0,B] and q in [0,D] (Assumption 3.2).
    This is the structural modeling premise that all covariate effects enter through the linear valuation index and the demand link is identical across contexts. The pilot-corrected residual representation and the entire rate depend on it.
  • domain assumption The induced link g(u) = E[q(epsilon_t - u)] is Holder-smooth on I_g with known beta and L_g (Assumption 3.4).
    The local-polynomial bias O(h^beta), the pilot-error correction, and the final rate all rely on this smoothness assumption. The paper shows it holds under smooth noise with bounded-variation q, but the theorem takes it as a primitive.
  • standard math Standard self-normalized martingale concentration, matrix determinant lemma, Bretagnolle-Huber inequality, and Holder Taylor remainder bounds.
    These are standard tools used throughout the appendix. They are not proved in the paper but are routine and correctly invoked.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue." pith.science (2026). https://pith.science/paper/WOEQIZ4Z

@misc{pith2026260803142,
  author       = {Pith},
  title        = {Pith review of: Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WOEQIZ4Z}},
  note         = {Machine review of arXiv:2608.03142}
}
read the original abstract

We study contextual dynamic pricing with arbitrary covariate sequences and bounded, possibly nonbinary purchase quantities. Demand follows a semiparametric surplus-index model with an unknown linear valuation parameter and an unknown H\"older-smooth response. We impose neither concavity nor strong unimodality on revenue and allow nonunique optimal prices. We develop a pilot-corrected layered decision-partitioning policy that combines directional pilot estimation, local polynomial learning, predictable data assignment, and global action elimination. Pilot correction removes the first-order effect of valuation-parameter error, while permanent labels enable concentration under adaptive sampling. The policy attains the minimax smoothness-dependent horizon rate up to logarithmic factors; a matching lower bound already holds for a constant-context binary-demand subclass.

Figures

Figures reproduced from arXiv: 2608.03142 by the authors.

Figure 1
Figure 1. Overview of the proposed adaptive semiparametric pricing policy. Panel (a) shows the uncertainty [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the lower-bound construction. Panel (a) constructs a smooth, nonincreasing baseline [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

70 extracted references · 38 canonical work pages

  1. [1]

    Operations Research , volume=

    Minimax Optimality in Contextual Dynamic Pricing with General Valuation Models , author=. Operations Research , volume=. 2025 , doi=

  2. [2]

    Forty-third International Conference on Machine Learning , year=

    The Cost of Information: Phase Transitions in Contextual Bandits with Paid Observations , author=. Forty-third International Conference on Machine Learning , year=

  3. [3]

    2009 , publisher=

    Discrete choice methods with simulation , author=. 2009 , publisher=

  4. [4]

    Econometrica: Journal of the Econometric Society , pages=

    Econometrics of first-price auctions , author=. Econometrica: Journal of the Econometric Society , pages=. 1995 , publisher=

  5. [5]

    2020 , publisher=

    Bandit algorithms , author=. 2020 , publisher=

  6. [6]

    Journal of Machine Learning Research , volume=

    Using confidence bounds for exploitation-exploration trade-offs , author=. Journal of Machine Learning Research , volume=

  7. [7]

    Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages=

    Contextual bandits with linear payoff functions , author=. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages=. 2011 , organization=

  8. [8]

    Advances in neural information processing systems , volume=

    Improved algorithms for linear stochastic bandits , author=. Advances in neural information processing systems , volume=

Show all 70 references
  1. [9]

    International Conference on Machine Learning , pages=

    Provably optimal algorithms for generalized linear contextual bandits , author=. International Conference on Machine Learning , pages=. 2017 , organization=

  2. [10]

    Mathematics of Operations Research , year=

    Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability , author=. Mathematics of Operations Research , year=

  3. [11]

    Advances in neural information processing systems , volume=

    A smoothed analysis of the greedy algorithm for the linear contextual bandit problem , author=. Advances in neural information processing systems , volume=

  4. [12]

    Artificial Intelligence and Statistics , pages=

    Contextual bandit learning with predictable rewards , author=. Artificial Intelligence and Statistics , pages=. 2012 , organization=

  5. [13]

    International Conference on Machine Learning , pages=

    A Reduction from Linear Contextual Bandit Lower Bounds to Estimation Lower Bounds , author=. International Conference on Machine Learning , pages=. 2022 , organization=

  6. [14]

    International Conference on Machine Learning , pages=

    Practical contextual bandits with regression oracles , author=. International Conference on Machine Learning , pages=. 2018 , organization=

  7. [15]

    International Conference on Machine Learning , pages=

    Beyond ucb: Optimal and efficient contextual bandits with regression oracles , author=. International Conference on Machine Learning , pages=. 2020 , organization=

  8. [16]

    Conference on Learning Theory , pages=

    Nearly minimax-optimal regret for linearly parameterized bandits , author=. Conference on Learning Theory , pages=. 2019 , organization=

  9. [17]

    arXiv preprint arXiv:2010.03104 , year=

    Instance-dependent complexity of contextual bandits and reinforcement learning: A disagreement-based perspective , author=. arXiv preprint arXiv:2010.03104 , year=

  10. [18]

    Journal of multivariate analysis , volume=

    The central limit theorem for weighted empirical processes indexed by sets , author=. Journal of multivariate analysis , volume=. 1987 , publisher=

  11. [19]

    International Conference on Machine Learning , pages=

    Sparsity-agnostic lasso bandit , author=. International Conference on Machine Learning , pages=. 2021 , organization=

  12. [20]

    International Conference on Artificial Intelligence and Statistics , pages=

    A parameter-free algorithm for misspecified linear contextual bandits , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2021 , organization=

  13. [21]

    , author=

    Minimax-Optimal Rates For Sparse Additive Models Over Kernel Classes Via Convex Programming. , author=. Journal of machine learning research , volume=

  14. [22]

    Proceedings of the National Academy of Sciences of the United States of America , volume=

    Minimax theorems , author=. Proceedings of the National Academy of Sciences of the United States of America , volume=. 1953 , publisher=

  15. [23]

    The annals of statistics , pages=

    Optimal global rates of convergence for nonparametric regression , author=. The annals of statistics , pages=. 1982 , publisher=

  16. [24]

    2009 , doi=

    Introduction to Nonparametric Estimation , author=. 2009 , doi=

  17. [25]

    2002 , publisher=

    A distribution-free theory of nonparametric regression , author=. 2002 , publisher=

  18. [26]

    The Annals of Statistics , volume=

    Regularization in kernel learning , author=. The Annals of Statistics , volume=. 2010 , publisher=

  19. [27]

    , author=

    Optimal Rates for Regularized Least Squares Regression. , author=. COLT , pages=

  20. [28]

    Foundations of Computational Mathematics , volume=

    Optimal rates for the regularized least-squares algorithm , author=. Foundations of Computational Mathematics , volume=. 2007 , publisher=

  21. [29]

    Conference on learning theory , pages=

    Beyond least-squares: Fast rates for regularized empirical risk minimization through self-concordance , author=. Conference on learning theory , pages=. 2019 , organization=

  22. [30]

    The Annals of Statistics , volume=

    Confidence bands in density estimation , author=. The Annals of Statistics , volume=. 2010 , doi=

  23. [31]

    The Fourteenth International Conference on Learning Representations , year=

    Semi-parametric contextual pricing with general smoothness , author=. The Fourteenth International Conference on Learning Representations , year=

  24. [32]

    arXiv preprint arXiv:2605.15411 , year=

    Harnessing Unimodality in Semiparametric Contextual Pricing via Oracle Price Map Learning , author=. arXiv preprint arXiv:2605.15411 , year=

  25. [33]

    arXiv preprint arXiv:2605.04207 , year=

    Optimal Semiparametric Dynamic Pricing with Feature Diversity , author=. arXiv preprint arXiv:2605.04207 , year=

  26. [34]

    arXiv preprint arXiv:2605.05609 , year=

    Optimal Contextual Pricing under Agnostic Non-Lipschitz Demand , author=. arXiv preprint arXiv:2605.05609 , year=

  27. [35]

    arXiv preprint arXiv:2405.06866 , year=

    Dynamic Contextual Pricing with Doubly Non-Parametric Random Utility Models , author=. arXiv preprint arXiv:2405.06866 , year=

  28. [36]

    International Conference on Artificial Intelligence and Statistics , pages=

    Incentive-aware contextual pricing with non-parametric market noise , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2023 , organization=

  29. [37]

    Mathematics of Operations Research , volume=

    Distribution-free contextual dynamic pricing , author=. Mathematics of Operations Research , volume=. 2024 , publisher=

  30. [38]

    Advances in Neural Information Processing Systems , volume=

    Contextual dynamic pricing with unknown noise: Explore-then-ucb strategy and improved regrets , author=. Advances in Neural Information Processing Systems , volume=

  31. [39]

    International Conference on Artificial Intelligence and Statistics , pages=

    Towards agnostic feature-based dynamic pricing: Linear policies vs linear valuation with unknown noise , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2022 , organization=

  32. [40]

    Journal of the American Statistical Association , volume=

    Policy optimization using semiparametric models for dynamic pricing , author=. Journal of the American Statistical Association , volume=. 2024 , publisher=

  33. [41]

    Journal of Machine Learning Research , volume=

    Dynamic pricing in high-dimensions , author=. Journal of Machine Learning Research , volume=

  34. [42]

    Advances in Neural Information Processing Systems , volume=

    Dynamic incentive-aware learning: Robust pricing in contextual auctions , author=. Advances in Neural Information Processing Systems , volume=

  35. [43]

    Algorithmic Learning Theory , pages=

    Dynamic pricing with finitely many unknown valuations , author=. Algorithmic Learning Theory , pages=. 2019 , organization=

  36. [44]

    International Conference on Machine Learning , pages=

    Semi-parametric contextual pricing algorithm using cox proportional hazards model , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  37. [45]

    arXiv preprint arXiv:2112.13254 , year=

    On dynamic pricing with covariates , author=. arXiv preprint arXiv:2112.13254 , year=

  38. [46]

    Operations Research , volume=

    Nonparametric pricing analytics with customer covariates , author=. Operations Research , volume=. 2021 , publisher=

  39. [47]

    arXiv preprint arXiv:2312.15999 , year=

    Pricing with Contextual Elasticity and Heteroscedastic Valuation , author=. arXiv preprint arXiv:2312.15999 , year=

  40. [48]

    Advances in Neural Information Processing Systems , volume=

    Improved Algorithms for Contextual Dynamic Pricing , author=. Advances in Neural Information Processing Systems , volume=

  41. [49]

    Available at SSRN 5133677 , year=

    Tight Regret Bounds in Contextual Pricing with Semi-parametric Demand Learning , author=. Available at SSRN 5133677 , year=

  42. [50]

    Surveys in operations research and management science , volume=

    Dynamic pricing and learning: historical origins, current research, and new directions , author=. Surveys in operations research and management science , volume=. 2015 , publisher=

  43. [51]

    Computer Communications , volume=

    Dynamic pricing techniques for Intelligent Transportation System in smart cities: A systematic review , author=. Computer Communications , volume=. 2020 , publisher=

  44. [52]

    Management Science , volume=

    Multimodal dynamic pricing , author=. Management Science , volume=. 2021 , publisher=

  45. [53]

    Management Science , volume=

    On the (surprising) sufficiency of linear models for dynamic pricing with demand learning , author=. Management Science , volume=. 2015 , publisher=

  46. [54]

    Advances in Neural Information Processing Systems , volume=

    Context-based dynamic pricing with partially linear demand model , author=. Advances in Neural Information Processing Systems , volume=

  47. [55]

    Management Science , year=

    Contextual Offline Demand Learning and Pricing with Separable Models , author=. Management Science , year=

  48. [56]

    The Annals of Statistics , volume=

    Transfer learning for contextual multi-armed bandits , author=. The Annals of Statistics , volume=. 2024 , publisher=

  49. [57]

    International Conference on Artificial Intelligence and Statistics , pages=

    Smoothness-Adaptive Dynamic Pricing with Nonparametric Demand Learning , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2024 , organization=

  50. [58]

    Operations Research , volume=

    Smoothness-adaptive contextual bandits , author=. Operations Research , volume=. 2022 , publisher=

  51. [59]

    Management Science , volume=

    Personalized dynamic pricing with machine learning: High-dimensional features and heterogeneous elasticity , author=. Management Science , volume=. 2021 , publisher=

  52. [60]

    Conference on Learning Theory , pages=

    Adaptivity to smoothness in x-armed bandits , author=. Conference on Learning Theory , pages=. 2018 , organization=

  53. [61]

    Operations research , volume=

    Dynamic pricing without knowing the demand function: Risk bounds and near-optimal algorithms , author=. Operations research , volume=. 2009 , publisher=

  54. [62]

    Management Science , volume=

    Dynamic pricing with external information and inventory constraint , author=. Management Science , volume=. 2024 , publisher=

  55. [63]

    Management Science , year=

    Context-based dynamic pricing with separable demand models , author=. Management Science , year=

  56. [64]

    Advances in Neural Information Processing Systems , volume=

    Online pricing for multi-user multi-item markets , author=. Advances in Neural Information Processing Systems , volume=

  57. [65]

    Mathematics of Operations Research , volume=

    A primal--dual learning algorithm for personalized dynamic pricing with an inventory constraint , author=. Mathematics of Operations Research , volume=. 2022 , publisher=

  58. [66]

    Available at SSRN 4803002 , year=

    LEGO: Optimal Online Learning under Sequential Price Competition , author=. Available at SSRN 4803002 , year=

  59. [67]

    Operations Research , year=

    To interfere or not to interfere: Information revelation and price-setting incentives in a multiagent learning environment , author=. Operations Research , year=

  60. [68]

    The Annals of Statistics , volume =

    Optimal Smoothing in Single-Index Models , author =. The Annals of Statistics , volume =

  61. [69]

    Semiparametric Least Squares (

    Ichimura, Hidehiko , journal =. Semiparametric Least Squares (

  62. [70]

    Journal of the American Statistical Association , volume =

    Direct Semiparametric Estimation of Single-Index Models with Discrete Covariates , author =. Journal of the American Statistical Association , volume =

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.