Pith. sign in

REVIEW 4 major objections 5 minor 86 references

Small open-weight LLMs, warm-started on optimal schedules and aligned with a dual-domain physics verifier, can produce jointly feasible electricity-and-compute schedules that cut grid violation and cost by an order of magnitude.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 23:20 UTC pith:YOCNPBZ6

load-bearing objection Solid systems paper: real joint ECCS task, oracle-backed ECBench, and SFT+physics-RL that actually moves 3B models—headline “beats frontier” claim is softer than the before/after gains. the 4 major comments →

arxiv 2607.26710 v1 pith:YOCNPBZ6 submitted 2026-07-29 cs.LG

PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems

classification cs.LG
keywords Electricity-Computing Co-SchedulingUnit CommitmentData-Center SchedulingLarge Language ModelsFeasibility-Aware RLECBenchDC Power Flow
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

AI data centers are becoming large, flexible, and volatile loads on the power grid. Scheduling them well means deciding both when generators run and where computing jobs land, under one set of physical and service constraints. General-purpose language models can write a schedule that looks valid, yet still violate line limits or leave load unserved. This paper defines that joint problem, builds a 2,000-instance benchmark with exact optimal solutions from a mixed-integer solver, and introduces PowerAtlas: an agent that first learns the format and patterns of optimal joint schedules, then is reinforced by a deterministic checker that scores both grid feasibility and task service. On three small open-weight models the method turns mostly unusable outputs into schedules that compete with much larger frontier models on violation and cost, in a single forward pass with no solver at inference. A provincial utility deployment validated the decision loop on real data-center data.

Core claim

PowerAtlas shows that supervised warm-start on Gurobi-optimal joint schedules, followed by Feasibility-Aware Group Relative Policy Optimization driven by a deterministic dual-domain verifier, is enough to make 3B-scale open-weight models emit electricity-computing co-schedules that are far more feasible and cheaper than the same models untrained, and competitive with frontier zero-shot APIs under the paper’s protocol, all in one solver-free forward pass.

What carries the argument

FA-GRPO (Feasibility-Aware Group Relative Policy Optimization): after full-parameter supervised fine-tuning on compliant optima, the policy is optimized with a group-relative objective whose trajectory reward first gates on parseable format, then scores compute service quality discounted by infeasibility and grid quality as oracle cost over agent cost plus valued lost load from a nodal DC feasibility LP.

Load-bearing premise

The reported gains rest on best-checkpoint selection and averages on the same held-out pool with no third split, and on frontier baselines scored mainly on one shared hard instance while open-weight rows use the full pool.

What would settle it

Re-run the three PowerAtlas backbones and the frontier baselines on the full 400-instance test pool under one fixed satisfaction convention, with checkpoints chosen on a true validation split never used for reporting; if the order-of-magnitude violation and cost gaps shrink or reverse, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Joint electricity-computing schedules can be produced at inference cost of one LLM forward pass rather than a mixed-integer solve that can take minutes in the tail.
  • Small open-weight models become usable operators’ aides for co-scheduling once they are warm-started on optima and graded by a physics verifier.
  • A public benchmark with oracle labels and a grid-aware evaluator becomes the standard yardstick for comparing ECCS methods.
  • Neither format learning nor reinforcement alone suffices: answerability from supervised init and feasibility from verifier-driven alignment must be stacked in that order.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same warm-start-plus-verifier pattern could transfer to other coupled infrastructure decisions where an exact solver exists offline but is too slow online, such as gas-power or water-energy co-dispatch.
  • Residual violation concentrated at evening peaks suggests the next lever is not better line routing but better peak supply or selective deferral of low-value compute.
  • If operators keep authority over grid actuation, the practical product is a fast proposal engine whose output is always re-checked, not a closed-loop controller.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper formalizes electricity-computing co-scheduling (ECCS) as a single LLM-agent decision that jointly produces 24-hour unit dispatch and cross-data-center task placement under DC power-flow, unit dynamics, rack capacity, and SLA constraints. It releases ECBench (2,000 RTS-GMLC + real data-center instances with Gurobi optima and a deterministic grid-aware evaluator) and PowerAtlas, a model-agnostic pipeline of full-parameter SFT on compliant optima followed by FA-GRPO whose terminal reward is built from a dual-domain verifier (format gate, Viol, cost efficiency vs c★, and compute satisfaction). On three 3B open-weight backbones, reported Viol falls by roughly 84–98% and Cost by up to ~26× relative to the untrained bases (Table 2); ablations attribute parseability to SFT and physics to FA-GRPO (Table 3). A provincial experimental network validates the decision loop functionally. Inference is one solver-free forward pass.

Significance. If the empirical protocol is tightened, this is a substantial systems+ML contribution: a coupled ECCS task with oracle-labeled instances, a reproducible physical evaluator, and evidence that small open-weight models can be aligned to hard multi-domain feasibility without a solver in the loop. Strengths that should be credited include the public code path, Gurobi-optimal labels, the independent hourly DC feasibility LP (Appendix A.4), train-only retrieval, honest residual analysis (mostly peak lost load), and the clear two-stage ablation. The work sits at a timely intersection of LLM agents and grid-aware data-center flexibility; the benchmark and verifier are reusable beyond the particular RL recipe.

major comments (4)
  1. [§6.1–6.2, Table 1–2, Appendix D.1] The headline claim that PowerAtlas agents 'match or beat the strongest frontier model' (§6.2, abstract) is not supported by a paired, same-protocol comparison. Appendix D.1 states that frontier rows in Table 1 are point estimates on one common high-penetration instance (m01d15-rho06-s10), while open-weight and PowerAtlas rows average the held-out pool; on that shared instance Llama+PowerAtlas Viol is 449.2 vs GPT-5.5's 445.8. Re-run all systems on the full held-out pool under identical decoding and report paired Viol/Cost (and confidence intervals), or restrict the competitiveness sentence to the single shared instance with all three backbones shown.
  2. [Appendix D.1, D.3; §6.1; Table 5] Checkpoint and stage selection use the same held-out pool that the tables report, with no third split (Appendix D.1, D.3). Shipped steps (FA-GRPO step 40; step 20 for Llama) are therefore best-checkpoint numbers on the evaluation set. This is load-bearing for the reported gains. Either freeze selection on a validation slice carved from train, or add a blind full-pool evaluation and mark current numbers as development-set results.
  3. [§4.2, §6, Tables 1–3, Appendix D.1] Satisfaction conventions are mixed across tables in a way that inflates the PowerAtlas narrative. D.1 states Table 1 frontiers (and some ablations) are grid-aware, while Table 2 PowerAtlas rows and several base/Figure 5 panels use placement; for untrained models the two differ by 2–3×, and Falcon3+PowerAtlas S-cnt moves from 69.5% (grid-aware) to 91.8% (placement). Viol and Cost are convention-invariant and should carry the main claim; every satisfaction column in the main tables must use one labeled convention, with the other in an appendix.
  4. [§4.1, §6.5, Fig. 5(e), Appendix D.3] Generalization is only seed-IID: train/test can share day and renewable penetration ρ (Appendix D.3). The robustness panel (Fig. 5e) still conditions within that split. A leave-one-month or leave-one-ρ evaluation is needed before claiming performance 'across operating conditions' (§6.5) under realistic regime shift; at minimum, report the seed-split limitation next to that claim.
minor comments (5)
  1. [§6.6, Abstract] Field deployment (RQ5, §6.6) is functional only; power-side actions are not actuated and savings are not reported. Soften abstract/intro language that may be read as operational grid impact.
  2. [Appendix C.1, Table 5] GRPO group size N=2 (Table 5, Appendix C) makes advantages noisy; briefly discuss sensitivity or note it as a limitation alongside the reward design that was introduced to compensate.
  3. [Fig. 5, §6.3, Appendix D.2] Figure 5(c)–(d) use training-side residual and placement rates that are not comparable to Tables 1–3; the caption says so, but a one-line reminder in §6.3 would help readers who skip the appendix.
  4. [§3, Appendix A.1] Notation: agent-controlled units are described as 6 including 1 swing in §3/A.1 but '5 adjustable non-swing units' in the decision description—keep one consistent phrasing.
  5. [§2] Related work is dense; a short explicit contrast table (single-domain LLM schedulers vs coupled ECCS with physical feedback) would clarify the novelty claim without lengthening the narrative.

Circularity Check

0 steps flagged

No derivation circularity: targets and rewards come from external Gurobi optima and an independent DC feasibility LP, not from the policy’s own outputs.

full rationale

PowerAtlas’s claimed chain is standard supervised-then-RL imitation of an external oracle, not a self-referential derivation. The ECCS program (Eq. 1) and the MILP in Appendix A.2 are solved by Gurobi to produce D★ and c★; SFT minimizes token CE to those oracle serializations (Eq. 7); FA-GRPO’s trajectory reward is gated by format then scored by a deterministic hourly nodal DC feasibility LP (Appendix A.4) that re-optimizes the swing unit and yields Viol, plus cost normalized to the same external c★ (Eqs. 11–14). None of these quantities is defined in terms of the learned policy’s predictions. Retrieval (Graph-R1/HyperGraphRAG) is train-split-only and explicitly unused in reported inference runs, so author-overlapping citations are not load-bearing for the feasibility/cost claims. Same-pool checkpoint selection and mixed placement vs grid-aware satisfaction conventions (Appendix D.1) are evaluation-validity issues, not definitional or fitted-input circularity under the stated patterns. The method is self-contained against external solver labels and physics checks; score 0 with empty steps.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 3 invented entities

The work is empirical systems/ML: claims rest on standard DC power-flow and UC modeling choices, a synthetic-but-calibrated task mix, Gurobi as oracle, VOLL and reward weights, and an evaluation protocol that equates reported gains with generalization. No new physical entities; free parameters are training/reward and scenario-generation knobs that shape the measured gains.

free parameters (6)
  • Reward weights (w_c, w_p)=(0.5,0.5), (η_v, η_c)=(0.7,0.3), φ_0=0.5, r_0=0.5 = 0.5/0.5, 0.7/0.3, 0.5, 0.5
    Hand-set FA-GRPO reward mix that defines the optimized objective; changes would re-rank feasibility vs service vs cost.
  • VOLL C_LOL = 9000 $/MWh
    Converts Viol into the grid quality term and dominates Cost; sets how strongly lost load drives both reward and economy metrics.
  • FA-GRPO steps / shipped checkpoint (step 40; step 20 for Llama) = shipped step 40 (20 for Llama-3.2-3B), 60 max
    Early-stopping point chosen on the held-out pool; Fig. 5(c) shows step 60 regresses—selection is a free experimental choice tied to reported numbers.
  • GRPO group size N=2, β=1e-3, clip ε=0.2, LR 2e-6 = N=2, β=1e-3, ε=0.2
    Advantage estimates from a single pair; authors note low within-group variance can kill gradient—training dynamics depend on these knobs.
  • Scenario grid: 50 days × ρ∈{1,2,4,6}% × 10 seeds; 37 tasks; 4 DCs × 11 racks = 2000 instances; tier mix ~50.8/21.9/27.3% H/M/L
    Defines ECBench difficulty and coupling strength; task distributions are synthesized and ‘deliberately relaxed in schedulability.’
  • Agent-controlled units = 6 (1 swing) of 73; spinning reserve γ=0.03 = 6 controllable; γ=0.03
    Restricts the decision space the LLM outputs; swing re-optimized in evaluator so Viol is residual imbalance after best swing action.
axioms (6)
  • domain assumption Linearized DC power flow with line limits and nodal balance is an adequate physics surrogate for scoring ECCS feasibility in this study.
    Used for oracle MILP and evaluator LP throughout §§3–4 and Appendix A; AC, contingency, and N-1 security are out of scope.
  • domain assumption Gurobi MILP optima are the correct reference decisions and costs for supervision and Q_grid normalization.
    Eq. (1), §4.1; median solve 16.8s—assumes model fidelity and solver optimality gaps are negligible for labeling.
  • ad hoc to paper IID split by generator seed measures relevant generalization for ECCS agents.
    §4.1 and D.3; authors admit test can share day and penetration with train—month/ρ hold-outs not used.
  • domain assumption Grid-aware evaluator (re-optimizing swing, slack-based Viol) is the right external judge of agent schedules.
    §4.2, A.4; underpins all Viol/Cost claims and FA-GRPO rewards.
  • standard math Token-level GRPO with terminal dual-domain reward is a valid policy gradient setup for structured schedule generation.
    §5.2 Eqs. (8)–(14); standard clipped importance ratio + group baseline + KL to ref.
  • domain assumption De-identified/synthesized task traces calibrated to one real DC generalize as ‘realistic physical operating conditions.’
    §4.1, §6.6; rack/power calibrated, task-level distributions synthesized.
invented entities (3)
  • ECBench independent evidence
    purpose: Standardized 2,000-instance ECCS benchmark with oracle labels and joint metrics.
    Constructed artifact, not a physical entity; value is empirical and depends on release fidelity.
  • FA-GRPO (Feasibility-Aware Group Relative Policy Optimization) no independent evidence
    purpose: Name for GRPO plus format gate and dual-domain quality reward using the physics verifier.
    Method label for a reward/shaping scheme; not an independent natural object.
  • PowerAtlas agent / ECCS task formalization no independent evidence
    purpose: Unified S→D LLM policy interface coupling D_P and D_C under F(S).
    Problem packaging and system name; substance is the training+verifier loop.

pith-pipeline@v1.2.0-daily-grok45 · 34955 in / 4544 out tokens · 96483 ms · 2026-07-30T23:20:09.175835+00:00 · methodology

0 comments
read the original abstract

The rapid growth of AI workloads is turning data centers into large-scale, volatile, yet spatiotemporally flexible grid loads, creating an urgent need for coordinated electricity-computing scheduling. Under stringent grid constraints, schedules from general-purpose large language models (LLMs) are often infeasible, causing line-flow violations and unserved load. We present PowerAtlas, an LLM-agent framework for electricity-computing co-scheduling that integrates historical instances, domain knowledge, and physical constraints to produce joint decisions satisfying both grid operational rules and the service-level agreements (SLAs) of computing tasks. Working with a provincial power utility in China, we built an experimental electricity-computing network and validated the decision loop on real data-center data; from de-identified operational data we further constructed ECBench, a benchmark of 2,000 scheduling instances with oracle-optimal solutions. Experiments across eleven LLMs demonstrate the effectiveness of PowerAtlas under realistic physical operating conditions, with consistent feasibility and cost gains across three open-weight backbones. Our code is publicly available at https://github.com/JAVA-Jiang/PowerAtlas.

Figures

Figures reproduced from arXiv: 2607.26710 by Anh Tuan Luu, Chao Yang, Haoran Luo, Kaiwen Jiang, Siya Xu, Ziyue Zhu.

Figure 1
Figure 1. Figure 1: An LLM agent maps a coupled specification to a [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview and motivation of PowerAtlas, illustrating the integration of large language models with power-system [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Task formulation of ECCS: an LLM agent maps an input specification [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The four-stage construction pipeline of ECBench: coupled scenario instances are synthesized from grid documentation, [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Results on ECBench. (a) Landscape of violation vs satisfaction (bubble size = cost). (b) Hourly violation and cumulative [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Failure cases on ECBench, quoted verbatim from the stored model answers. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Operations dashboard of the co-scheduling system [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: The four templates that define the PowerAtlas interface: the instance serialization and output contract (a), the [PITH_FULL_IMAGE:figures/full_fig_p016_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Four answers to one ECBench instance (m01d15-rho06-s10), quoted verbatim from the stored model outputs, with the optimum for reference. Each row is one agent-controlled unit over hours 1–24, in percent of its operating range; off is shutdown and 0 is minimum stable output [PITH_FULL_IMAGE:figures/full_fig_p017_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

86 extracted references · 12 canonical work pages · 6 internal anchors

  1. [1]

    Henrik Abgaryan, Ararat Harutyunyan, and Tristan Cazenave. 2024. LLMs can Schedule. arXiv preprint. arXiv:2408.06993. doi:10.48550/arXiv.2408.06993

  2. [2]

    Ankit Aharwar, Ram Naresh, Veena Sharma, and Vineet Kumar. 2023. Unit Commitment Problem for Transmission System, Models and Approaches: A Review.Electric Power Systems Research223 (2023), 109671. doi:10.1016/j.epsr. 2023.109671

  3. [3]

    Ali AhmadiTeshnizi, Wenzhi Gao, and Madeleine Udell. 2024. OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235). PMLR, 577–596. https://proceedings. mlr.press/v235/ahmaditeshnizi24a.html

  4. [4]

    Ahsan Ali and Öznur Özkasap. 2024. Spatial and Thermal Aware Methods for Efficient Workload Management in Distributed Data Centers.Future Generation Computer Systems153 (2024), 360–374. doi:10.1016/j.future.2023.12.006

  5. [5]

    Clayton Barrows, Aaron Bloom, Ali Ehlen, Jussi Ikaheimo, Jennie Jorgenson, Dheepak Krishnamurthy, Jessica Lau, Brendan McBennett, Matthew O’Connell, Eugene Preston, Andrea Staid, Gord Stephen, and Jean-Paul Watson. 2019. The IEEE Reliability Test System: A Proposed 2019 Update.IEEE Transactions on Power Systems35, 1 (2019), 119–127. doi:10.1109/TPWRS.2019.2925557

  6. [6]

    Fabien Bernier, Jun Cao, Maxime Cordy, and Salah Ghamizi. 2025. PowerGraph- LLM: Novel Power Grid Graph Embedding and Optimization with Large Lan- guage Models.IEEE Transactions on Power Systems40, 6 (2025), 5483–5486. doi:10.1109/TPWRS.2025.3596774

  7. [7]

    Yifan Bian, Lirong Xie, Lan Ma, and Hangong Zhang. 2024. A Novel Two-Stage Energy Sharing Method for Data Center Cluster Considering Carbon-Green Certificate Coupling Mechanism.Energy313 (2024), 133991. doi:10.1016/j.energy. 2024.133991

  8. [8]

    Boyu Chen, Yanbo Che, Zhihao Zheng, and Shuaijun Zhao. 2023. Multi-Objective Robust Optimal Bidding Strategy for a Data Center Operator Based on Bi-Level Optimization.Energy269 (2023), 126761. doi:10.1016/j.energy.2023.126761

  9. [9]

    Kaixuan Chen, Wei Luo, Shunyu Liu, Yaoquan Wei, Yihe Zhou, Yunpeng Qing, Quan Zhang, Yong Wang, Jie Song, and Mingli Song. 2025. Powerformer: A Section-Adaptive Transformer for Power Flow Adjustment. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1. Association for Computing Machinery, Toronto, ON, Canada, 2204–22...

  10. [10]

    Xavier, Tongxin Zheng, Muham- mad Marwali, Bernard Knueven, Yongpei Guan, Peter B

    Yonghong Chen, Feng Pan, Feng Qiu, Alinson S. Xavier, Tongxin Zheng, Muham- mad Marwali, Bernard Knueven, Yongpei Guan, Peter B. Luh, Lei Wu, Bing Yan, Mikhail A. Bragin, Haiwang Zhong, Anthony Giacomoni, Ross Baldick, Boris Gisin, Qun Gu, Russ Philbrick, and Fangxing Li. 2022. Security-Constrained Unit Commitment for Electricity Market: Modeling, Solutio...

  11. [12]

    Yuheng Cheng, Huan Zhao, Xiyuan Zhou, Junhua Zhao, Yuji Cao, Chao Yang, and Xinlei Cai. 2025. A Large Language Model for Advanced Power Dispatch. Scientific Reports15 (2025), 8925. doi:10.1038/s41598-025-91940-x

  12. [14]

    Yingrui Fan and Junbo Zhao. 2026. Harnessing Flexible Spatial and Temporal Data Center Workloads for Grid Regulation Services. arXiv preprint. doi:10. 48550/ARXIV.2602.01508

  13. [15]

    Xiaomin Fang, Jizhou Huang, Fan Wang, Lingke Zeng, Haijin Liang, and Haifeng Wang. 2020. ConSTGAT: Contextual Spatial-Temporal Graph Attention Network for Travel Time Estimation at Baidu Maps. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. Association for Computing Machinery, 2697–2705. doi:10.1145/3394...

  14. [16]

    Zheng Fang, Qingqing Long, Guojie Song, and Kunqing Xie. 2021. Spatial- Temporal Graph ODE Networks for Traffic Flow Forecasting. InProceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. Association for Computing Machinery, 364–373. doi:10.1145/3447548.3467430

  15. [17]

    Yang Fu, Xiaoyan Guo, Yang Mi, Minghan Yuan, Xiaolin Ge, Xiangjing Su, and Zhenkun Li. 2021. The distributed economic dispatch of smart grid based on deep reinforcement learning.IET Generation, Transmission & Distribution15, 18 (2021), 2645–2658. doi:10.1049/gtd2.12206

  16. [18]

    Fang Gao, Guojian Wu, Suhang Guo, Wei Dai, and Feng Shuang. 2023. Solving DC Power Flow Problems Using Quantum and Hybrid Algorithms.Applied Soft Computing137 (2023), 110147. doi:10.1016/j.asoc.2023.110147

  17. [19]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schel- ten, Alex Vaughan, et al. 2024. The Llama 3 Herd of Models. arXiv preprint. arXiv:2407.21783. doi:10.48550/arXiv.2407.21783

  18. [20]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, et al. 2025. DeepSeek-R1 Incentivizes Reasoning in LLMs through Reinforcement Learning.Nature645, 8081 (2025), 633–638. doi:10.1038/s41586-025-09422-z

  19. [21]

    Gurobi Optimization, LLC. 2024. Gurobi Optimizer Reference Manual. https: //www.gurobi.com

  20. [22]

    Ouzhu Han, Tao Ding, Miao Yang, Wenhao Jia, Xinran He, and Zhoujun Ma

  21. [23]

    Songqiao Han, Xiyang Hu, Hailiang Huang, Minqi Jiang, and Yue Zhao. 2022. ADBench: Anomaly Detection Benchmark. InAdvances in Neural Informa- tion Processing Systems, Vol. 35. Curran Associates, Inc., New Orleans, LA, USA, 32142–32159. https://proceedings.neurips.cc/paper_files/paper/2022/hash/ cf93972b116ca5268827d575f2cc226b-Abstract-Datasets_and_Benchm...

  22. [24]

    Dengshan Hou, Li Wang, Yanru Ma, Longbiao Lyu, Weijie Liu, and Shenghu Li. 2025. Joint Optimal Scheduling of Power Grid and Internet Data Centers Considering Time-of-Use Electricity Price and Adjustable Tasks for Renewable Power Integration.Sustainability17, 8 (2025), 3374. doi:10.3390/su17083374

  23. [25]

    Chenyu Huang, Zhengyang Tang, Shixi Hu, Ruoqing Jiang, Xin Zheng, Dongdong Ge, Benyou Wang, and Zizhuo Wang. 2025. ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling.Operations Research73, 6 (2025), 2986–3009. doi:10.1287/opre.2024.1233

  24. [26]

    2025.Energy and AI

    International Energy Agency. 2025.Energy and AI. Technical Report. IEA, Paris. https://www.iea.org/reports/energy-and-ai

  25. [27]

    Prachi Jadhav, Hongwei Jin, Ewa Deelman, and Prasanna Balaprakash. 2025. Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling. InProceedings of the SC’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. 2234–2244

  26. [28]

    Mengshuo Jia, Zeyu Cui, and Gabriela Hug. 2025. Enhancing LLMs for Power Sys- tem Simulations: A Feedback-driven Multi-agent Framework.IEEE Transactions on Smart Grid(2025)

  27. [29]

    Yongqing Jiang, Jianze Wang, Zhiqi Shen, Zhenghong Lin, Jiayuan Wang, Yi- jian Yang, Kaoshan Dai, and Haoran Luo. 2026. Rethinking Scientific Modeling: Toward Physically Consistent and Simulation-Executable Programmatic Genera- tion. arXiv preprint. arXiv:2602.07083. doi:10.48550/arXiv.2602.07083

  28. [30]

    Abdullahi Bala Kunya, Adamu Saidu Abubakar, and Samuel Sunday Yusuf. 2023. Review of Economic Dispatch in Multi-Area Power System: State-of-the-Art and Future Prospective.Electric Power Systems Research217 (2023), 109089. doi:10.1016/j.epsr.2022.109089

  29. [32]

    Hunter Lightman, Vineet Kosaraju, Yura Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe

  30. [33]

    Lesieutre, and Line Roald

    Julia Lindberg, Yasmine Abdennadher, Jiaqi Chen, Bernard C. Lesieutre, and Line Roald. 2021. A Guide to Reducing Carbon Emissions through Data Center Geographical Load Shifting. InProceedings of the Twelfth ACM International Conference on Future Energy Systems. Association for Computing Machinery, Virtual Event, Italy, 430–436. doi:10.1145/3447555.3466582

  31. [34]

    InInternational Conference on Learn- ing Representations

    Let’s Verify Step by Step. InInternational Conference on Learn- ing Representations. https://proceedings.iclr.cc/paper_files/paper/2024/hash/ aca97732e30bcf1303bc22ac3924fd16-Abstract-Conference.html

  32. [35]

    Shaohui Liu, Sungho Shin, and Deepjyoti Deka. 2026. Watts vs. Bytes: Turning Data Centers into Grid Assets via Storage Compute Co-Optimization. arXiv preprint. arXiv:2605.16190. doi:10.48550/arXiv.2605.16190

  33. [36]

    Lesieutre, and Line A

    Julia Lindberg, Bernard C. Lesieutre, and Line A. Roald. 2022. Using Geographic Load Shifting to Reduce Carbon Emissions.Electric Power Systems Research212 (2022), 108586. doi:10.1016/j.epsr.2022.108586

  34. [37]

    Haoran Luo, Haihong E, Guanting Chen, Qika Lin, Yikai Guo, Fangzhi Xu, Zemin Kuang, Meina Song, Xiaobao Wu, Yifan Zhu, and Anh Tuan Luu. 2026. Graph- R1: Towards Agentic GraphRAG Framework via End-to-End Reinforcement Learning. InProceedings of the 43rd International Conference on Machine Learning. https://openreview.net/forum?id=YXnFGsSkCC

  35. [38]

    Xiaoou Liu. 2024. Research on Collaborative Scheduling of Internet Data Cen- ter and Regional Integrated Energy System Based on Electricity-Heat-Water Coupling.Energy292 (2024), 130462. doi:10.1016/j.energy.2024.130462

  36. [39]

    Haoran Luo, Haihong E, Yikai Guo, Qika Lin, Xiaobao Wu, Xinyu Mu, Wen- hao Liu, Meina Song, Yifan Zhu, and Anh Tuan Luu. 2025. KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search. InProceed- ings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267). PMLR, Vancouver, Canad...

  37. [40]

    Haoran Luo, Haihong E, Guanting Chen, Yandan Zheng, Xiaobao Wu, Yikai Guo, Qika Lin, Yu Feng, Zemin Kuang, Meina Song, Yifan Zhu, and Anh Tuan Luu. 2025. HyperGraphRAG: Retrieval-Augmented Generation via Hypergraph-Structured Knowledge Representation. InAdvances in Neural Information Processing Systems, Vol. 38. Curran Associates, Inc., San Diego, CA, USA...

  38. [41]

    Haoxiang Luo, Kun Yang, Qi Huang, Marco Aiello, and Schahram Dustdar. 2025. A Novel Hierarchical Co-Optimization Framework for Coordinated Task Sched- uling and Power Dispatch in Computing Power Networks. arXiv preprint. doi:10.48550/ARXIV.2508.04015

  39. [42]

    Haoran Luo, Haihong E, Zichen Tang, Shiyao Peng, Yikai Guo, Wentai Zhang, Chenghao Ma, Guanting Dong, Meina Song, Wei Lin, Yifan Zhu, and Anh Tuan Luu. 2024. ChatKBQA: A Generate-then-Retrieve Framework for Knowledge Base Question Answering with Fine-tuned Large Language Models. InFindings of the Association for Computational Linguistics: ACL 2024. Associ...

  40. [45]

    Sina Mohammadi, Ali Hassan, Rouzbeh Haghighi, Van-Hai Bui, and Wencong Su. 2025. Large Language Models for Solving Economic Dispatch Problem. In 2025 IEEE Energy Conversion Congress and Exposition (ECCE). IEEE, 1–5

  41. [46]

    Chuizheng Meng, Sirisha Rambhatla, and Yan Liu. 2021. Cross-Node Feder- ated Graph Neural Network for Spatio-Temporal Data Modeling. InProceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. Association for Computing Machinery, 1202–1211. doi:10.1145/3447548.3467371

  42. [47]

    Ozgur Kayalica, Denizhan Guven, A

    Beltus Wiysobunri Nkwawir, M. Ozgur Kayalica, Denizhan Guven, A. Can Du- man, and Hamza Salih Erden. 2025. Carbon-Aware Workload Management in Data Centers: A Multi-Energy Integration Approach. InProceedings of the 16th ACM International Conference on Future and Sustainable Energy Systems. Association for Computing Machinery, Rotterdam, The Netherlands, 9...

  43. [48]

    Varma, Hang Zou, Qiyang Zhao, and Merouane Debbah

    Thomas Mongaillard, Samson Lasaulce, Othman Hicheur, Chao Zhang, Lina Bariah, Vineeth S. Varma, Hang Zou, Qiyang Zhao, and Merouane Debbah. 2024. Large Language Models for Power Scheduling: A User-Centric Approach. In Proceedings of the 22nd International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks. IEEE, Seoul, Republi...

  44. [49]

    Zheyi Pan, Yuxuan Liang, Weifeng Wang, Yong Yu, Yu Zheng, and Junbo Zhang

  45. [50]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training Language Models to Follow Instructions with Hum...

  46. [51]

    Jiaju Qi, Lei Lei, Kan Zheng, and Simon X. Yang. 2022. Joint Energy Dispatch and Unit Commitment in Microgrids Based on Deep Reinforcement Learning. arXiv preprint. arXiv:2206.01663. doi:10.48550/arXiv.2206.01663

  47. [52]

    Ana Radovanović, Ross Koningstein, Ian Schneider, Bokan Chen, Alexandre Duarte, Binz Roy, Diyue Xiao, Maya Haridasan, Patrick Hung, Nick Care, Saurav Talukdar, Eric Mullen, Kendal Smith, MariEllen Cottman, and Walfredo Cirne

  48. [53]

    Dekang Qi, Xiuwen Yi, Chengjie Guo, Yanyong Huang, Junbo Zhang, Tianrui Li, and Yu Zheng. 2024. Spatio-Temporal Consistency Enhanced Differential Network for Interpretable Indoor Temperature Prediction. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery, 5590–5601. doi:10.1145/363752...

  49. [54]

    Muhammad Sarwar, Muhammad Rizwan, Mubushra Aziz, and Abdul Rehman Sudais. 2025. Large Language Models for Power System Applications: A Com- prehensive Literature Survey. arXiv preprint. arXiv:2512.13004. doi:10.48550/ arXiv.2512.13004

  50. [55]

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language Models Can Teach Themselves to Use Tools. InAdvances in Neural Information Processing Systems, Vol. 36. Curran Associates, Inc. https: //openreview.net/forum?id=Yacmpz84TH

  51. [56]

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov

  52. [57]

    Jiaqi Ruan, Gaoqi Liang, Huan Zhao, Guolong Liu, Xianzhuo Sun, Jing Qiu, Zhao Xu, Fushuan Wen, and Zhao Yang Dong. 2024. Applying Large Language Models to Power Systems: Potential Security Threats.IEEE Transactions on Smart Grid 15, 3 (2024), 3333–3336. doi:10.1109/TSG.2024.3373256

  53. [58]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Yu Wu, and Daya Guo. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.arXiv preprint arXiv:2402.03300(2024)

  54. [59]

    Zezhi Shao, Zhao Zhang, Fei Wang, and Yongjun Xu. 2022. Pre-Training En- hanced Spatial-Temporal Graph Neural Network for Multivariate Time Series Forecasting. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery, 1567–1577. doi:10.1145/3534678.3539396

  55. [60]

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning. InAdvances in Neural Information Processing Systems, Vol. 36. Curran Associates, Inc., 8634–8652. https://openreview.net/forum?id=vAElhFcKW6

  56. [61]

    Chengguo Su, Lingshuang Wang, Quan Sui, and Huijun Wu. 2025. Optimal Scheduling of a Cascade Hydro-Thermal-Wind Power System Integrating Data Centers and Considering the Spatiotemporal Asynchronous Transfer of Energy Resources.Applied Energy377 (2025), 124360. doi:10.1016/j.apenergy.2024.124360

  57. [62]

    Rana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang, Minlan Yu, and Michael Mitzenmacher. 2025. Don’t Stop Me Now: Embedding-Based Sched- uling for LLMs. InInternational Conference on Learning Representations (ICLR). 63345–63368. https://proceedings.iclr.cc/paper_files/paper/2025/file/ 9eb8b5ccb0de594a16548f7c058fdadf-Paper-Conference.pdf

  58. [63]

    Technology Innovation Institute. 2024. The Falcon 3 Family of Open Models. Technical blog post. https://falcon-lm.github.io/blog/falcon-3/

  59. [64]

    Haoxiang Wan and Xingpeng Li. 2026. Data Center Spatio-Temporal Load Flexi- bility in Security-Constrained Unit Commitment for Enhanced Grid Efficiency and Reliability. arXiv preprint. arXiv:2605.18517. doi:10.48550/arXiv.2605.18517

  60. [65]

    Haoyu Wang, Hongke Guo, Zhaoliang Zhu, You Zhang, Yu Zhou, and Xudong Zheng. 2024. BacktrackSTL: Ultra-Fast Online Seasonal-Trend Decomposition with Backtrack Technique. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery. doi:10.1145/3637528.3671510

  61. [66]

    Jiangjiang Wang, Hongda Deng, Yi Liu, Zeqing Guo, and Yongzhen Wang. 2023. Coordinated Optimal Scheduling of Integrated Energy System for Data Center Based on Computing Load Shifting.Energy267 (2023), 126585. doi:10.1016/j. energy.2022.126585

  62. [67]

    Babak Taheri and Daniel K. Molzahn. 2024. Optimizing Parameters of the DC Power Flow.Electric Power Systems Research235 (2024), 110719. doi:10.1016/j. epsr.2024.110719

  63. [68]

    Qingsong Wen, Zhe Zhang, Yan Li, and Liang Sun. 2020. Fast RobustSTL: Effi- cient and Robust Seasonal-Trend Decomposition for Time Series with Complex Patterns. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. Association for Computing Machinery, 2203–2213. doi:10.1145/3394486.3403271

  64. [69]

    Philipp Wiesner, Ilja Behnke, Dominik Scheinert, Kordian Gontarska, and Lauritz Thamsen. 2021. Let’s Wait Awhile: How Temporal Workload Shifting Can Reduce Carbon Emissions in the Cloud. InProceedings of the 22nd International Middleware Conference. Association for Computing Machinery, 260–272. doi:10. 1145/3464298.3493399

  65. [70]

    Chris Williams, Philip Colangelo, Ayse Coskun, Ethan Levine, Andy Neale, Ciaran Roberts, Shayan Sengupta, Nikhil Shirolkar, Varun Sivaram, Sarah Soares, Ethan Tiao, Scott Underwood, Daniel Wilson, Frank Sharp, Luke Wainwright, Harry Petty, Scott Wallace, and Brandon Records. 2026. Power-Flexible AI Data Centers: A New Paradigm for Grid-Responsive Compute....

  66. [71]

    Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, Xiaojun Chang, and Chengqi Zhang. 2020. Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural Networks. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. Association for Computing Machinery, 753–763. doi:10.1145/3394486.3403118 PowerAt...

  67. [72]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of- Thought Prompting Elicits Reasoning in Large Language Models. InAdvances in Neural Information Processing Systems, Vol. 35. Curran Associates, Inc., 24824–24837. https://proceedings.neurips.cc/paper_files/paper/2022/hash/ 9...

  68. [73]

    Ziyang Xiao, Dongxiang Zhang, Yangjun Wu, Lilin Xu, Yuan Wang, Xiong- wei Han, Xiaojin Fu, Tao Zhong, Jia Zeng, Mingli Song, and Gang Chen. 2024. Chain-of-Experts: When LLMs Meet Complex Operations Research Problems. In International Conference on Learning Representations. OpenReview.net, Vienna, Austria, 48519–48537. https://proceedings.iclr.cc/paper_fil...

  69. [74]

    Qian Xu, Chutian Yu, Xiang Yuan, Zao Fu, and Hongzhe Liu. 2023. A Privacy- Preserving Distributed Subgradient Algorithm for the Economic Dispatch Prob- lem in Smart Grid.IEEE/CAA Journal of Automatica Sinica10, 7 (2023), 1625–1627. doi:10.1109/JAS.2022.106028

  70. [75]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 Technical Report. arXiv preprint. arXiv:2505.09388. doi:10.48550/arXiv.2505.09388

  71. [76]

    Le, Denny Zhou, and Xinyun Chen

    Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen. 2024. Large Language Models as Optimizers. InInterna- tional Conference on Learning Representations. https://openreview.net/forum? id=Bb4VGOWELI

  72. [77]

    Rogier Hans Wuijts, Marjan van den Akker, and Machteld van den Broek. 2024. Effect of Modelling Choices in the Unit Commitment Problem.Energy Systems 15, 1 (2024), 1–63. doi:10.1007/s12667-023-00564-5

  73. [78]

    Xu Yang, Chenhui Lin, Yue Yang, Qi Wang, Haotian Liu, Haizhou Hua, and Wenchuan Wu. 2026. Large Language Model Powered Automated Modeling and Optimization of Active Distribution Network Dispatch Problems.IEEE Transactions on Smart Grid17, 2 (2026), 952–965. doi:10.1109/TSG.2025.3621438

  74. [79]

    Griffiths, Yuan Cao, and Karthik Narasimhan

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. InAdvances in Neural Information Processing Systems, Vol. 36. Curran Associates, Inc. https://openreview.net/forum?id= 5Xc1ecxO1h

  75. [80]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Mod- els. InInternational Conference on Learning Representations. https://openreview. net/forum?id=WE_vluYUL-X

  76. [81]

    Junchen Ye, Zihan Liu, Bowen Du, Leilei Sun, Weimiao Li, Yanjie Fu, and Hui Xiong. 2022. Learning the Evolutionary and Multi-Scale Graph Structure for Multivariate Time Series Forecasting. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery, 2296–2306. doi:10.1145/3534678.3539274

  77. [82]

    Linxiao Yang, Rui Ren, Xinyue Gu, and Liang Sun. 2023. Interactive Generalized Additive Model and Its Applications in Electric Load Forecasting. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery. doi:10.1145/3580305.3599848

  78. [83]

    Paschalidis, and Ayse K

    Yijia Zhang, Athanasios Tsiligkaridis, Ioannis Ch. Paschalidis, and Ayse K. Coskun. 2024. Data Center and Load Aggregator Coordination Towards Electric- ity Demand Response.Sustainable Computing: Informatics and Systems42 (2024), 100957. doi:10.1016/j.suscom.2024.100957

  79. [84]

    Yuanshi Zhang, Bokang Zou, Xu Jin, Yifu Luo, Meng Song, Yujian Ye, Qinran Hu, Qirui Chen, and Antonio Carlos Zambroni. 2025. Mitigating power grid impact from proactive data center workload shifts: A coordinated scheduling strategy integrating synergistic traffic - data - power networks.Applied Energy 377 (2025), 124697. doi:10.1016/j.apenergy.2024.124697

  80. [85]

    Haoruo Zhao, Mathieu Tanneau, and Pascal Van Hentenryck. 2022. A Linear Outer Approximation of Line Losses for DC-Based Optimal Power Flow Problems. Electric Power Systems Research212 (2022), 108272. doi:10.1016/j.epsr.2022.108272

Showing first 80 references.