{"id":"b6496a5b-181d-4647-97cf-4d1b7d6c4dca","arxiv_id":"2608.00921","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A Thompson-sampling Whittle-index framework for data center demand response is extended with gated priors, low-rank smoothing, and adaptive mixing, and is shown in simulations to outperform baseline policies in sparse-data stress tests.","lead":"This paper proposes a learning-based system, RACER+, that lets power grids ask data centers to cut load without seeing their internal job schedules. It combines restless bandit algorithms with Thompson sampling and shows in simulations that refined variants beat baseline policies under stress conditions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5.2's 'every refined variant beats TW' is not supported by Table 3, which reports only the per-cell best refinement; post-hoc cellwise selection undermines the robustness claim.","rationale":"The reader's CONDITIONAL verdict is appropriate. The most load-bearing threat to the central empirical claim is not the Markov-state assumption (which, if violated, affects the oracle and all compared policies similarly and is at least partially addressed by the contextual features x_i(t) and the supplementary shared-queue assumption), but the evaluation protocol: the headline 'Every refined variant beats TW' is not what Table 3 shows. The table reports only the best refined variant per cell, selected after the fact. This makes the 15–34 point margins suspect as estimates of the advantage one would obtain by deploying any particular refinement. I do not think this warrants REJECT, because the narrow 'cellwise best' claim in the abstract is directly supported by the table's margins vs TW, and the methodological variants are well-specified. But the paper's stronger prose and the 'pre-registered' spectral rule need either the full per-variant results (supplementary Section 6) or a fixed-variant stress sweep. The proposed test settles whether the concern lands. I partially agree with the reader: they also flag post-hoc selection and the count contradiction in their rationale, though their stated weakest assumption is the Markov state representation, which I regard as secondary.","tokens_in":11376,"tokens_out":5511,"duration_ms":51063,"concrete_test":"Re-run the 16-cell stress sweep with one fixed refined variant—e.g., Adaptive TW + beta-gated prior, or the spectral rank rule applied uniformly to every cell—and report margins versus TW and versus the strongest baseline for all cells. If the fixed variant does not beat TW in every cell at p<0.05, the robustness claim is a post-hoc selection artifact; the abstract should then be corrected to report only the per-cell best.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that refined Thompson-Whittle variants are robustly superior depends on Table 3, but Table 3 reports only the best refined variant in each (|S|, ρ) cell, with the identity of that variant changing across cells (e.g., Adp. TW + beta gated prior, Adp. TW + support/offline prior, Adp. TW + beta support-gated prior + low-rank). The 15–34 point margins versus TW are therefore maxima over a family of refinements chosen after observing the outcomes. Choosing the best variant per cell inflates the apparent gain, and the paper gives no pre-specified rule for selecting a refined variant in a new cell. The spectral rank rule described in Section 5.3 is presented as 'pre-registered' but is introduced after the fact with a p-value from the same 16-cell sweep. Section 5.2 is also internally inconsistent: it first says the best refined variant 'leads the strongest baseline in fourteen, tying in two whose intervals span zero' and then says it exceeds the strongest baseline in 15 of 16 settings. Because the abstract's claim is deliberately limited to 'cellwise best', the table does support that narrow statement, but the paper's stronger prose ('Every refined variant beats TW in all sixteen cells') and the practical implication of a deployable robust policy do not follow.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes RACER+, a contextual restless multi-armed bandit (RMAB) framework for data-center demand-response flexibility scheduling. Each data center is modeled as an MDP whose state is the current position in a cyclic job queue; transition and reward functions are learned online via Thompson sampling, and arm selection is made by a Whittle index policy computed from posterior samples. The paper contributes several refinements to the basic Thompson–Whittle (TW) learner: an adaptive UCB/TW mixer, offline and support priors, beta-gated priors, low-rank transition smoothing, and a spectral rank rule. Experiments cover a baseline regime, a transition-stress regime with |S| in {8, 20, 50, 100} and noise probability rho in {0, 0.1, 0.2, 0.3} over 100 rounds, and a real-world Texas data-center simulation. The headline claim is that, in every stress cell, the cellwise best refined variant beats TW by 15–34 points of oracle reward and also beats the strongest baseline in most cells, at lower computational cost than EXP4.","tokens_in":11604,"tokens_out":9950,"duration_ms":79778,"significance":"The application is timely and the empirical setup is more realistic than a purely synthetic bandit benchmark: the authors use Azure VM and MIT SuperCloud traces, provide an anonymized code supplement, and report seed-level error bars that permit paired comparisons. The refined variants address a genuine sparse-data difficulty, and the L1/off-support diagnostics are a sensible way to attribute gains. If a single pre-specified refined variant could be shown to dominate TW and strong baselines across the stress sweep, this would be a useful advance for demand-response scheduling without operator visibility. At present, however, the central robustness claim is established only for a per-cell post-hoc selected variant, so the practical significance is not yet demonstrated; the revision should either provide a fixed selection rule or reframe the claims.","major_comments":[{"comment":"The sentence 'Every refined variant beats TW in all sixteen cells by 15–34 points of oracle' is not supported by Table 3, which reports only the single best refined variant per (|S|,rho) cell, with the winning variant changing across cells (e.g., 'Adp. TW + beta gated prior' at (8,0.0) and 'Adp. TW' at (100,0.1)). Because the per-cell winner is selected after observing the outcomes, the reported margins are maxima over a family of refinements and do not demonstrate that any single refined policy robustly beats TW. The abstract's 'cellwise best' wording is appropriately limited, but the prose and the practical framing of a deployable robust policy are not. Please report all variants or a pre-specified selection rule, and adjust the claims.","section":"Section 5.2, Table 3"},{"comment":"The text is internally inconsistent: it first says the best refined variant 'leads the strongest baseline in fourteen, tying in two whose intervals span zero' and then says it 'surpasses the average reward of the strongest baseline in 15 out of 16 settings.' Table 3 shows negative margins versus the baseline at (|S|=8,rho=0.2) (-0.71 +/- 6.27) and (|S|=100,rho=0.3) (-2.90 +/- 13.87), i.e., two losses, not two ties, so neither count matches the table. Please correct the counts and the interpretation.","section":"Section 5.2"},{"comment":"The spectral rank rule is described as 'pre-registered,' but it is introduced after the results and evaluated on the same 16-cell stress sweep used to develop it; the reported gain (+2.13 points, 95% CI [+1.55,+2.72], p<10^-5, n=120) is a post-hoc comparison on the same data, not a confirmatory test. Please provide evidence that the rule was fixed before observing these cells (e.g., a dated protocol) or re-frame the spectral analysis as exploratory with a separate hold-out evaluation.","section":"Section 5.3"},{"comment":"The per-cell paired t-tests reported in Table 3 are unadjusted for multiple comparisons across 16 cells and numerous refined variants; with this many tests, p<0.05 is expected under the global null. Reporting family-wise adjusted p-values (e.g., Holm-Bonferroni) or pre-specifying a single primary cell/variant comparison would substantially strengthen the robustness claim.","section":"Section 5.2, Table 3"},{"comment":"The MDP state is defined only as 'positions i in the circular job queue' in Table 1, but the reward from rescheduling (power savings minus delay penalty) and the queue evolution depend on which specific jobs are in the look-ahead window, which is not part of the state. If the queue-position state is not sufficient for the transition and reward functions, the learned P and R are misspecified and the Whittle indices computed from them do not correspond to the true control problem. The text mentions an 'additional assumption' in the supplementary shared-queue formulation but does not validate it in the main text. Please add a validation (e.g., predictive checks of the learned transition/reward model, or a comparison with a state-enriched variant).","section":"Table 1 and Equations (2), (5)"},{"comment":"The text says 'We also perform the hyperparameter sweeps to identify optimal values, such as the trust floor tau_min in (0,0.1] and beta gate G_i^a.' If these parameters are tuned on the same stress cells for which Table 3 reports improvements, the reported margins are optimistically biased. Please specify which hyperparameters were fixed before the experiments and evaluate sensitivity on held-out settings (e.g., nested resampling), or temper the claims accordingly.","section":"Section 4.1"}],"minor_comments":[{"comment":"'placed the increasing pressures' should be 'placing increasing pressure', and 'Lamma 1' should be 'Lemma 1'.","section":"Sections 1 and 3.1"},{"comment":"'The gives a combination of across 16 runs' is ungrammatical; please revise.","section":"Section 5.2"},{"comment":"The caption uses 'refined TM-TW variants' although the text defines the variants as TW; please harmonize the acronyms.","section":"Figure 2 caption"},{"comment":"'The EXP4 strategy constantly achieves the best performance in the refinement family' is confusing because EXP4 is a baseline, not a refinement; please rephrase (e.g., 'EXP4 is the strongest baseline').","section":"Section 5.2"},{"comment":"The citation 'Dai et al., Liu et al., 2026' is incomplete; supply full author lists or citation keys.","section":"Section 2"},{"comment":"There are minor typographical issues: 'Whittle Whittle [1988]' duplicates the author name, 'Micheal Terrall' should be 'Michael Terrall', and 'p-test' in Table 3 should be 'p-value' or 'paired t-test p-value'.","section":"References and Table 3"},{"comment":"The abstract claims 'lower computational cost than EXP4,' but no runtime measurements or complexity analysis appear in the main text; please add evidence or qualify the claim.","section":"Abstract and Section 5.2"},{"comment":"Several theoretical claims (indexability in Definition 1, the L1/off-support leakage analyses in Section 5.1, and the low-rank composed-kernel certificate in Section 3.3) are deferred to supplementary material that was not supplied with the manuscript; please ensure the supplement is available at review time so these claims can be checked.","section":"Supplementary material"}],"recommendation":"major_revision","confidential_remarks":"The paper's practical claims are stronger than the per-cell-best analysis supports. The editorial decision should hinge on whether the authors can provide a pre-specified selection rule and corrected multiple-comparison statistics. I also note that the supplementary material and code were not available for review; they should be provided with the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi [Name],\n\nQuick take: this is a useful applied paper, not a breakthrough. It takes the known Thompson-sampled Whittle-index machinery for restless bandits and adapts it to VM-level scheduling for grid demand response, with a set of pragmatic refinements—gated priors, low-rank transition smoothing, support priors, adaptive mixing with UCB—and evaluates them on Azure and MIT SuperCloud traces. That combination is new and probably of real interest to the data-center/grid community. The calibration effort, the economic framing, and the honesty about needing dataset-specific tuning are all points in its favor. The citation pattern is fine; it builds on Whittle 1988, Chen and Hou 2024, and related bandit work, with no red flags.\n\nWhere it falls short is the evidence for the central claim. The paper says every refined variant beats TW in all sixteen stress cells, but Table 3 actually reports only the per-cell best refined variant, and the identity of that variant jumps across cells. So the 15–34 point margins are maxima over a family of candidates chosen after seeing the outcomes, not evidence for any single deployable policy. Within the table itself there is also a small credibility problem: the text mentions fourteen wins and two 'ties' whose intervals span zero, but those two cells have negative point margins (−0.71 and −2.90), not ties; later the text says 15 of 16. The p-values are unadjusted and the 'pre-registered' spectral rule for rank selection appears in the results section with p-values computed from the same 16-cell sweep. I'd want a fixed evaluation protocol: pick one variant specification before the sweep, or report the full distribution of all variants, not just the best.\n\nThe modeling also leans on an unexamined Markov assumption: state is the queue position, but the reward and next state depend on which jobs are actually in the lookahead window, not just the position. The paper defers to the supplementary shared-queue formulation but doesn't validate sufficiency. This could matter for the Whittle indices.\n\nNet: the framework is plausible and the refinements make sense as heuristics, but the paper currently overstates its robustness. If the supplement and code were available and the evaluation protocol fixed, it would be a decent venue-level paper. As posted, it deserves a serious referee but needs revision before acceptance.\n\nBest.","headline":"Useful applied extension of Thompson-Whittle RMAB to data center demand response, but the headline robustness claim outruns the reported evidence because the best variant is selected per stress cell and the supplement that would settle it is absent.","tokens_in":12186,"tokens_out":2843,"would_cite":false,"duration_ms":26199,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that prior-augmented Thompson–Whittle scheduling lets data centers answer grid flexibility requests without revealing job details, and that refined variants beat plain Thompson–Whittle in every reported stress test.","keywords":["restless multi-armed bandits","Whittle index","Thompson sampling","data center demand response","job rescheduling","Markov decision process","low-rank smoothing","VM workload traces"],"falsifier":"Run the learned policy on a job trace in which the same queue position regularly yields very different rescheduling power savings, such as mixing huge AI training jobs with tiny interactive jobs, and compare the estimated transition and reward functions against the empirical frequencies; if the queue-position model cannot predict realized savings, the Whittle indices are computed from mis-specified models and the 15–34 point margins should shrink or reverse.","tokens_in":11130,"feed_emoji":"⚡","tokens_out":7302,"duration_ms":59776,"temperature":0.7,"pith_summary":"This paper tries to show that a data center can offer power-grid flexibility services by rescheduling batches of virtual-machine jobs, without ever revealing internal job details, and that the scheduling can be learned online from coarse observations. It builds each data center's cyclic job queue as a small Markov decision process, estimates the unknown transition and reward functions with Thompson sampling, and derives Whittle-index policies from the sampled estimates. Because Thompson–Whittle alone struggles when the state space is large and visits are sparse, the paper adds several refinements: an adaptive mix of the Whittle score with UCB exploration, gated structural priors, low-rank transition smoothing, and offline/support priors. In simulated baseline and stress tests on real VM traces, the authors report that every refined variant beats plain Thompson–Whittle in all sixteen stress cells by 15–34 points of oracle reward, and that the best refined variant beats the strongest baseline in fourteen cells and ties in two. A reader should care because this is a concrete route to making data centers economic, auditable demand-response participants under realistic privacy constraints.","feed_headline":"Refined bandit scheduling beats classic policy in all 16 stress tests","feed_subtitle":"Domain priors and adaptive mixing lift learned rescheduling to 70-96% of oracle reward under sparse, noisy conditions.","key_machinery":"The load-bearing machinery is the Whittle index computed from posterior samples of each arm's transition and reward functions, together with refinements that discipline those samples. The Whittle index is the smallest subsidy that makes leaving an arm passive optimal, which decouples the multi-armed problem into per-arm single-agent MDPs and yields a rankable score. The refinements are: an adaptive trust weight $\\tau_i(s,t)$, which blends the sampled Whittle score with a UCB score in states that have few observed transitions; a $\\beta$-gated prior that replaces a sampled transition row by a convex combination of the posterior row and a structural queue-based prior, with the gate drawn from a Beta distribution to avoid over-conservatism; low-rank SVD smoothing of transition matrices to pool information across rows; and offline/support priors that encode feasible transitions via zero masks and historical queue models. These mechanisms act on the estimates that feed the Whittle calculation rather than on the final action directly, so they preserve the index-policy interpretation.","core_discovery":"The central discovery is that prior-augmented Thompson sampling makes Whittle-index scheduling usable in sparse, noisy data-center environments, where the plain Thompson–Whittle policy degrades sharply. On the paper's own account, the adaptive mixed strategy that weights the sampled Whittle index against a UCB score by visit frequency, combined with domain priors (gated structural priors and offline/support masks) and low-rank smoothing of transition rows, lifts cumulative reward from roughly 41.8–61.1% of the oracle for plain TW to 65.6–96.0% for the best refined variant across the reported stress sweep. The real-world simulation on eight data centers moves from 41.9% of oracle for adaptive TW to 78.5% with the full refinement. The paper also establishes that the improved reward is accompanied by lower $\\ell^1$ transition-estimation error and less probability mass on infeasible transitions, though at 1,000 rounds the best transition estimator is not the highest reward earner. The authors conclude that domain knowledge, not just more data, is what lets a learned Whittle policy survive sparse state visits.","pith_inferences":["The paper's state omits which jobs are in the lookahead window, so the reported margins may shrink in settings where identical queue positions hide very different rescheduling savings; testing on job-level traces with explicit power models would settle this.","The spectral rank rule found in the appendix—choosing the smallest rank that retains 95% of singular-value mass—looks like a general fix for low-rank transition smoothing, not one tied to this application.","The adaptive trust weight could be transferred to other restless-bandit learning problems as a generic exploration–exploitation device, since it only needs visit counts and a sampled index.","The economic gains quoted in dollars per kWh suggest a concrete market test: a grid operator could compare realized load reductions under this policy against a rule-based baseline during peak price windows."],"forward_implications":["Grid operators can request load reductions through a learned policy that never needs to see job internals; the data center only reports coarse context and a flexibility action.","In sparse regimes—short learning horizons, large state spaces, noisy state observations—the refined variants recover a large share of the oracle reward where plain TW collapses.","Because the refined family is cheaper than EXP4's expert mixing, the improvement is not bought with extra computation.","The demonstrations on real VM traces suggest the approach transfers from lightweight internal workloads to ML training and inference jobs.","If the Markov model holds, the same index-policy machinery could be reused across data centers without retraining the full joint problem."],"supporting_citations":[{"why":"Supplies the index-policy definition and Lagrangian relaxation that the whole scheduling approach is built on.","marker":"[Whittle, 1988]"},{"why":"Provides the Microsoft Azure VM dataset used in the baseline and transition-stress experiments.","marker":"[Cortez et al., 2017]"},{"why":"Provides the calibrated real-world ML/AI workload dataset for the eight-data-center simulation.","marker":"[The MIT Supercloud, 2021]"},{"why":"Supplies the EXP4 expert-mixing baseline that the refined variants are compared against.","marker":"[Beygelzimer et al., 2011]"},{"why":"Grounds the Dirichlet–Categorical Bayesian prior used for transition estimation.","marker":"[Hjort, 2010]"},{"why":"Supplies the low-rank SVD approximation used to smooth sparse transition matrices.","marker":"[Jaderberg et al., 2014]"},{"why":"Establishes the contextual RMAB demand-response setting that the paper extends.","marker":"[Chen and Hou, 2024]"}],"fun_headline_variants":["Prior-augmented Thompson beats plain Whittle in sparse bandits","Domain priors rescue Whittle scheduling under sparse data","Refined bandit hits 96% of oracle in stress tests","Learned Whittle policy survives sparse data with domain priors","Thompson+prior beats vanilla Whittle in adaptive bandits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the current position in the job queue is a sufficient state, so the power savings and the next position can be predicted from it alone; in practice the specific jobs waiting in the lookahead window also determine those outcomes.","fun_headline_variants_meta":{"raw":{"variants":["Prior-augmented Thompson beats plain Whittle in sparse bandits","Domain priors rescue Whittle scheduling under sparse data","Refined bandit hits 96% of oracle in stress tests","Learned Whittle policy survives sparse data with domain priors","Thompson+prior beats vanilla Whittle in adaptive bandits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000554,"raw_usage":{"total_tokens":2672,"prompt_tokens":1012,"completion_tokens":1660,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":1575}},"tokens_in":628,"tokens_out":1660,"duration_ms":10862,"temperature":1.0,"reasoning_tokens":1575,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:14:55.644015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the learned policy on a job trace in which the same queue position regularly yields very different rescheduling power savings, such as mixing huge AI training jobs with tiny interactive jobs, and compare the estimated transition and reward functions against the empirical frequencies; if the queue-position model cannot predict realized savings, the Whittle indices are computed from mis-specified models and the 15–34 point margins should shrink or reverse.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the calibrated real-world ML/AI workload dataset for the eight-data-center simulation."}],"review_version":1}