{"id":"dd086182-deaa-4036-9ef8-b559fcb64fc5","arxiv_id":"2607.05024","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"A minimal directional perturbation of the max-difference revenue statistic yields an asymptotically valid p-value for structural properties of the oracle assortment under high-dimensional sparse contextual MNL with adaptive collection.","lead":"The paper builds a p-value for whether a learned optimal product assortment satisfies a structural rule (e.g., includes core items or meets category mix) after adaptive high-dimensional learning. It matters because platforms need graded evidence on operational constraints without conservative uniform error bounds over huge candidate sets.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Assumption 4.4 is the load-bearing sufficient condition for local Hessian stability; its anti-concentration step controls the Cn rates that underwrite the debiased expansion and Theorems 4.5–4.6.","rationale":"The reader correctly isolates Assumption 4.4 as the weakest link. After line-by-line inspection of the validity proof (S.2.5), the martingale coupling (Lemma S.2.2), the localization to ¯S0\\cup¯S1, and the directional covering argument are all standard once the rates are in hand; the only place those rates can break is the anti-concentration step that produces Cn in Claim S.3.1. The simulation design itself violates both branches of Assumption 4.4 (\\rho is typically hundreds, and full SK has large intersections), yet empirical size still controls for large T; this is consistent with the authors’ “sufficient not necessary” remark but confirms that the theorems do not yet cover the numerical evidence. No deeper internal inconsistency appears, the localization and \\\\chi^{2} calibration are appropriately conservative for uniform validity over arbitrary boundary geometries, and the regret bound is of the expected order. Hence the CONDITIONAL verdict with low correctness risk remains appropriate; the concrete check above would simply quantify how much the proof can be relaxed.","tokens_in":52661,"tokens_out":846,"duration_ms":40765,"concrete_test":"Under the exact simulation DGP of Section 5 (full SK, K=3, n=20, features N(0,I_p/3) truncated, revenues N(6.5,1)), draw 500 independent (v,\\beta) pairs near \\beta* and compute the empirical max pairwise correlation \\gamma of {\\xi_S} together with the realized max choice-probability ratio \\rho. If 1-\\gamma\\ll K^{-2} or \\rho\\gg2 (as the feature/coefficient scales suggest), re-derive the Lipschitz constant of \\beta\\mapsto E[\\Sigma_t(\\beta)|\\beta] on a fine grid; if it exceeds the Cn of Assumption 4.4 by more than a constant factor, the proof rates do not cover the experiments and the validity claim needs a weaker local-stability hypothesis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central size/power claims (Theorems 4.5–4.6) rest on the debiased expansion of Corollary 4.4 and the uniform revenue linearization of Lemma S.2.3, both of which inherit the estimation rates of Theorem 4.1. Those rates in turn require the local continuity bound ||E{\\Sigma_t(\\beta1)|\\beta1}-E{\\Sigma_t(\\beta2)|\\beta2}||_max \\lesssim Cn \\nu^{2}\\|\\beta1-\\beta2\\|_{1} + \\nu^{2}/T. Claim S.3.1 obtains this bound from the total-variation distance between the adaptive selection distributions P(St(\\beta)|\\beta,vt) and P(St(\\beta*)|\\beta*,vt). The TV distance is controlled by anti-concentration of the difference of maxima of the Gaussian assortment-revenue process {\\xi_S} (Theorem 2.4 of the cited work [5]), which needs either (i) the full K-subset class together with \\rho\\le2 (to guarantee 1-\\gamma\\gtrsim K^{-2}) or (ii) the restricted-intersection condition on SK. If either fails, Cn can become large enough that the T-thresholds of Theorem 4.5 are no longer sufficient for the remainders to be absorbed by \\kappa, and the Gaussian-coupling error \\eta of Lemma S.2.2 likewise fails to be o(\\sqrt s*). Remark 4.1 correctly notes that the assumption is only sufficient, yet the theorems as stated stand or fall with it.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper develops post-learning inference for structural properties of combinatorial optimizers in a high-dimensional sparse contextual MNL assortment model under adaptive data collection. The target is whether the terminal oracle optimizer intersects a prescribed structural class S0 (product inclusion, category proportion, feature screens). The authors reformulate this as a sign test on a nonsmooth max-difference revenue functional and propose a minimal directional perturbation p-value: after ℓ1-penalized online likelihood estimation and debiasing on the selected support, random unit directions capture angular uncertainty while the minimal radius UT needed to reach the null boundary is calibrated via a χ²_bs tail plus a directional residual δm. Theory provides uniform rates (Thm 4.1), effective support recovery (Cor 4.2), a debiased martingale expansion (Cor 4.4), and asymptotic size/power under localization to near-boundary assortments (Thms 4.5–4.6), using anti-concentration for Gaussian maxima differences and martingale Gaussian coupling. Simulations show size control at large T and substantially higher power than a uniform-error-bound baseline across three structural tests, with sublinear regret for the adaptive policy.","tokens_in":53191,"tokens_out":1506,"duration_ms":19518,"significance":"If the results hold, the paper supplies a usable inferential tool for a practically important and statistically irregular problem: testing discrete structural properties of data-dependent combinatorial optimizers after adaptive high-dimensional learning, where the parameter-to-optimizer map is discontinuous and standard Wald/delta-method tools fail. The localization of uncertainty to the max-difference boundary (rather than uniform control over SK) is a clear conceptual and power advantage over confidence-set inversion. The technical apparatus—anti-concentration for differences of maxima under adaptive selection, martingale coupling of the adaptive score, and effective-support debiased expansions—is of independent interest for irregular post-selection and post-adaptive-sampling inference. Explicit non-asymptotic rates, full validity/power proofs, and a regret guarantee that embeds inference in online learning without a pure exploration phase are genuine strengths. The contribution is well positioned at the intersection of high-dimensional inference, combinatorial optimization, and sequential decision-making.","major_comments":[{"comment":"Assumption 4.4 (Gaussian revenues plus either full K-subsets with ρ≤2 or a restricted-intersection condition on SK) is load-bearing for the local Hessian stability bound that drives the Cn rates in Lemma S.2.1, Claim S.3.1, the debiased expansion (Cor 4.4), and the T-thresholds of Theorems 4.5–4.6. Remark 4.1 correctly flags sufficiency and sketches a Stein-kernel relaxation, but the main theorems are not proved under any weaker primitive. For a paper that advertises a general max-difference perturbation principle for combinatorial optimizers, the manuscript should either (i) state the theorems under an abstract local-Hessian-stability condition with Ass 4.4 as a corollary, or (ii) supply a complete proof under a clearly weaker revenue/tail condition. As written, the scope of the validity claim is narrower than the introduction suggests.","section":"Assumption 4.4, Remark 4.1, Claim S.3.1"},{"comment":"The theoretically sufficient order κ ≍ σ_{vT,rT} √(s*/(Tλ)) depends on the unknown local gradient scale σ_{vT,rT} and λ_min(Σ*). The implementable choice κ = Cκ √(bs/(Tϵ)) with Cκ = 10^{-4} works in the reported simulations, but Theorem 4.5’s remainder absorption (display (S.2.35), (S.2.40), (S.2.42)) requires κ to dominate several higher-order terms involving Cn, ν, and ηT. The paper needs a clearer finite-sample prescription or a data-driven rule for Cκ (or for estimating σ_{vT,rT} on the localized sets S̄0, S̄1) so that practitioners can verify the conditions under which the o(1) size guarantee is expected to kick in.","section":"Theorem 4.5, Remark 4.6, Section 5.1"},{"comment":"Size is evaluated only at the least-favorable boundary Δ* = 0 obtained by bisection on a single terminal revenue coordinate (Section 5.1). This is informative for Type I control at the knife-edge, but it does not address size under interior nulls (Δ* > 0) or under the stronger null S*_T ⊆ S0 of Remark 2.1. At least one additional size panel under unmodified null contexts with Δ* > 0, and a brief discussion of how the procedure would be modified for the strong null, would make the empirical size claim more complete.","section":"Section 5.1, Remark 2.1"}],"minor_comments":[{"comment":"Notation for the selected support size switches between bs and s* after support recovery; a single convention after Corollary 4.2 would reduce cognitive load.","section":"Section 3.2, Theorem 4.5"},{"comment":"Figures 1–3 truncate the size axis at 0.10; a short note in the caption that early-horizon size can exceed this range (as already stated in the text) would prevent misreading of the plots.","section":"Figures 1–3"},{"comment":"The directional residual δm is defined with ϵ in (16), but the simulation protocol fixes δm = 0.002 and backs out m (capped at 50,000). State explicitly whether the cap ever binds for s ∈ {3,4,5} and what is done if it does.","section":"Section 5.1, Eq. (16)"},{"comment":"Several references to working manuscripts ([6], [5]) are central to the anti-concentration and comparison arguments; ensure arXiv or published versions are cited if available at revision.","section":"References"},{"comment":"Typographical: “COMBINA TORIAL” and “INFORMA TION” in the title header appear to have spurious spaces; “PENGYULI” / “SHUTINGSHEN” spacing in the author line should be cleaned.","section":"Title page"}],"recommendation":"minor_revision","confidential_remarks":"The technical core looks carefully executed and the problem is genuinely interesting for math.ST / sequential decision audiences. The main risk is overselling generality relative to Assumption 4.4; if the authors reframe the theorems around an abstract Hessian-stability condition, the paper is close to ready. Fit for a top statistics journal is good; less so for a pure OR venue without more algorithmic content."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the real thing for post-learning inference on combinatorial optimizers. The target is not β or a smooth functional but whether the oracle terminal assortment sits in a structural class (inclusion, category mix, feature screen). They reduce it to the sign of a nonsmooth max-difference of revenues and calibrate a p-value by the minimal radius of random unit-sphere perturbations on the debiased selected support. Directional uncertainty is handled by sphere sampling; magnitude by a χ^{2} tail. That localization is the practical payoff: they avoid the conservative uniform revenue-error bound over the whole SK.\n\nWhat is new is the combination: online ℓ1-penalized collection with sublinear regret, effective-support recovery, debiased expansion, martingale Gaussian coupling for the adaptive score, and anti-concentration for differences of Gaussian maxima to control selection-induced Hessian variation. Theorems 4.1, 4.5–4.6 and the corollaries are written carefully; the proofs are complete and checkable. Simulations on three concrete tests show size settling and substantially higher power than the UEB baseline, especially near the boundary (Example 2). Citations are honest about the gaps relative to Shen et al. and the low-dim contextual work.\n\nThe soft spot is Assumption 4.4 (Gaussian revenues plus either full K-subsets with ρ≤2 or restricted intersections). It supplies the local Hessian stability and the anti-concentration that drive the Cn rates. Remark 4.1 correctly calls it sufficient and sketches Stein-kernel and sub-Gaussian relaxations, but the theorems as stated stand on it. Finite-sample tuning of Cλ, ϵ, Cκ, δm is also needed and there is no public code. Neither sinks the contribution.\n\nThis is for people who do high-dimensional adaptive inference, assortment/revenue management, or post-selection for discrete decisions. It deserves a serious referee. I would engage with it and expect to cite the construction.","headline":"Solid, usable theory for a real irregular problem: max-difference directional perturbation after adaptive high-dim assortment learning, with complete proofs and clear power gains over uniform calibration.","tokens_in":53711,"tokens_out":530,"would_cite":true,"duration_ms":6846,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62F07","62J12","62J07","62L10"],"pacs":[],"model":"grok-4.5","headline":"A minimal directional perturbation of the max revenue gap yields a valid p-value for whether a learned assortment optimizer satisfies a structural constraint.","keywords":["irregular inference","high-dimensional inference","adaptive data collection","directional perturbation test","post-regularization inference","combinatorial optimization","contextual multinomial logit","assortment optimization"],"falsifier":"In the reported simulation design (n=20, p=500, K=3, s*=3–5, least-favorable null with revenue gap exactly zero), if the empirical Type-I rate of the proposed p-value stays systematically above 0.05 as T grows to 2000 while the uniform-error baseline remains near zero, the asymptotic size claim fails in the regime the theory targets.","tokens_in":53570,"feed_emoji":"📊","tokens_out":749,"duration_ms":6151,"temperature":0.7,"pith_summary":"After adaptive high-dimensional learning of customer preferences, a platform may care less about the exact best product assortment than about whether some optimal assortment still meets a business rule: include core products, keep category balance, or obey an inventory screen. Because the map from preferences to the best assortment jumps at ties, ordinary confidence intervals do not work. This paper reduces the structural question to the sign of a single nonsmooth number—the gap between the best revenue inside the rule class and the best revenue outside it—and builds a p-value by asking how large a random directional nudge of the estimated revenue surface is needed before that gap becomes compatible with the null. The data are gathered by an online sparse likelihood policy that also keeps cumulative revenue loss sublinear. Under a localized signal condition the resulting p-value controls size and has power, while avoiding the conservatism of calibrating uniform error over every feasible assortment.","feed_headline":"Perturb the revenue gap, not every assortment","feed_subtitle":"A minimal-radius directional test gives valid p-values for structural rules after adaptive assortment learning","key_machinery":"The minimal directional perturbation radius UT: the smallest a ≥ 0 such that, along at least one of m random unit directions on the selected support, the perturbed max-difference of null versus alternative plug-in revenues exceeds -κ; the p-value is the χ^{2} upper tail of UT^{2} plus a directional discretization remainder.","core_discovery":"For a high-dimensional contextual multinomial logit model under adaptive assortment selection, the structural hypothesis that the terminal oracle optimizer intersects a prescribed combinatorial class is equivalent to non-negativity of a max-difference revenue functional. A p-value formed from the minimal radius of random unit-sphere perturbations of the debiased terminal revenue surface on the selected support is asymptotically valid under the null and consistent under a localized alternative, without uniform error control over the full candidate class.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Minimal directional radius tests structural optimizer classes","Revenue-gap min-perturbation gives post-adaptive p-values","Localize inference near the null with unit-sphere revenue noise","Support-restricted max-difference for combinatorial class tests","Debiased revenue surface perturbations after adaptive assortment"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"The main rates and validity proofs rest on a primitive condition that revenues are Gaussian and the feasible assortment class is either the full K-subset class with bounded choice-probability ratios or a restricted-intersection subclass; that condition supplies the local Hessian stability and anti-concentration used throughout.","fun_headline_variants_meta":{"raw":{"variants":["Minimal directional radius tests structural optimizer classes","Revenue-gap min-perturbation gives post-adaptive p-values","Localize inference near the null with unit-sphere revenue noise","Support-restricted max-difference for combinatorial class tests","Debiased revenue surface perturbations after adaptive assortment"]},"model":"grok-4.5","effort":"low","cost_usd":0.006088,"raw_usage":{"total_tokens":1548,"prompt_tokens":800,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":60880000,"prompt_tokens_details":{"text_tokens":800,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":668,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":800,"tokens_out":80,"duration_ms":5677,"temperature":1.0,"reasoning_tokens":668,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T09:54:54.321929+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"In the reported simulation design (n=20, p=500, K=3, s*=3–5, least-favorable null with revenue gap exactly zero), if the empirical Type-I rate of the proposed p-value stays systematically above 0.05 as T grows to 2000 while the uniform-error baseline remains near zero, the asymptotic size claim fails in the regime the theory targets.","supporting_citations":[],"review_version":1}