{"id":"d63e4d04-1236-4086-8856-1422d6c6cc93","arxiv_id":"2506.06359","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A systematic review of 99 papers on transformers and LLMs in energy, ending with a proposal for LLM-powered Agentic Digital Twins for smart grids.","lead":"This paper systematically reviews 99 studies on transformers and large language models in energy forecasting and grid management, and proposes that these models will turn digital twins into 'Agentic Digital Twins'. It is a useful map of a fast-growing field for energy researchers and grid operators deciding where to invest.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Agentic DT claim extrapolates from offline LLM assistants in building modeling to autonomous grid management, without evidence that LLM-generated actions are safe or feasible; the paper's own hallucination and human-in-loop caveats undercut the 'autonomous' and 'proactive' attributes.","rationale":"This pass evaluates the paper's central claim in Sec. 7: LLMs will turn digital twins into Agentic Digital Twins that are autonomous, proactive, communicative, and capable of planning and recommending energy management actions. For that claim to be sound, LLM-based agents must exhibit reliable situational awareness and decision support in operational, safety-critical energy contexts. The reviewed evidence does not establish this. The 99-paper corpus includes LLM applications in carbon emission prediction, electrical safety incident detection, wind/load forecasting, building energy modeling (BEM), IDF generation, retrofit recommendation, and one chatbot/grid-management support paper ([30]); none demonstrates an LLM-driven agent running live grid operations, and none measures the safety or feasibility of LLM-suggested actions. The paper is transparent about this: Sec. 6 lists hallucination risk, lack of action validation, and the need for simulation-engine checks, and it explicitly states that LLMs cannot enforce actions and that humans remain in the loop. These admissions are exactly the conditions under which the 'autonomous' and 'proactive' attributes of an Agentic DT would fail. Thus the central claim is an extrapolation from offline, human-supervised modeling tasks to real-time multi-agent coordination.\n\nA secondary but related weakness is corpus selection: the citation-count cutoffs (Table 1) and the restriction to WoS plus four publishers exclude many 2024-2025 LLM-agent and arXiv works, several of which the authors themselves cite in the perspectives section (e.g., [125], [127], [131], [135]) but did not include in Table 2. This inconsistency makes it impossible for the review to claim that LLM-agent work in energy is 'early' or that the Agentic DT concept is novel; the field may already contain stronger evidence. The reader's weakest_assumption (selection bias) captures this second issue; we agree partially, but the more load-bearing problem is the logical gap between the reviewed tasks and the central claim.\n\nWhat would settle it: a small prototype as described in concrete_test. If an LLM planner hooked to a microgrid simulator cannot produce reliably feasible, safe recommendations without frequent human correction, then the 'autonomous, proactive' characterization is not supported, and the paper's value rests on its review and on a clearly labeled research agenda. That is consistent with the reader's conditional acceptance; no verdict change is needed, but the condition should require the authors to relabel Sec. 6-7 as speculative vision and to correct the [103] misclassification and the corpus/arXiv inconsistency.","tokens_in":36113,"tokens_out":5825,"duration_ms":60983,"concrete_test":"Prototype a minimal Agentic DT: connect an LLM planner (e.g., GPT-4o with retrieval-augmented access to grid constraints) to a simulated campus microgrid (GridLAB-D or OpenAI Gym), run 30 simulated days of hour-ahead operation, and have a domain simulator automatically verify each recommended action for feasibility and safety. Measure the rate of infeasible or unsafe actions and the number of human overrides, and compare against a rule-based or optimization baseline. If the LLM agent produces a high rate of infeasible actions or requires constant operator correction, the 'autonomous, proactive' central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Sec. 7) is that LLMs will transform digital twins into Agentic Digital Twins that are 'autonomous, proactive, and communicative' and capable of 'planning and recommending energy management actions.' The evidence marshalled for this is (i) a handful of LLM applications, mostly in building energy modeling (BEM), code/IDF generation, retrofit recommendation, and load forecasting, and (ii) discussions of LLM-enabled RAG, multimodal data fusion, and action planning. None of the 99 reviewed papers demonstrates an LLM-driven digital twin operating in a real-time grid or microgrid context, and none measures the reliability of LLM-generated management actions. The paper itself lists open challenges: hallucination risk, lack of action validation, need for simulation-engine checks, and the assertion that 'LLMs cannot execute or enforce actions directly in the smart grid' and 'humans will remain a key in the decision-making loop.' Consequently, the step from 'LLMs can assist in offline modeling tasks' to 'DTs will become autonomous agents that can act, communicate, and negotiate' is a large extrapolation. The load-bearing assumption is that capabilities shown in non-safety-critical, single-building, human-in-the-loop settings will transfer to safety-critical, multi-asset, real-time coordination with acceptable reliability. The review provides no evidence for this transfer; it is a vision, not a synthesized finding. Additionally, because the corpus selection excludes low-citation 2024-2025 work and non-WoS venues (e.g., arXiv preprints such as [125], [127], [131], [135], which are cited in the perspectives but not in Table 2), the review cannot confirm that this extrapolation is needed—the field may already contain more advanced LLM-agent/grid-control studies than the filtered corpus reveals. Thus the central claim is simultaneously under-supported by the reviewed evidence and potentially outdated by the excluded literature.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a PRISMA-guided systematic review of 99 papers (2021-2025, Web of Science, four publishers) applying transformer models and LLMs in the energy sector. It synthesizes transformer-based forecasting of electricity, thermal, wind, solar, and EV loads; reviews LLM applications in building energy modeling, scenario generation, anomaly detection, and decision support; and proposes the Agentic Digital Twin concept, in which LLM integration gives digital twins autonomy, proactivity, and social interaction. The review is a qualitative synthesis rather than a meta-analysis; no fitted equations or code are included. The paper's main empirical conclusions are that transformers dominate forecasting, LLMs extend capabilities beyond prediction, and early multi-agent LLM workflows are emerging; the Agentic DT is presented as the next step.","tokens_in":36417,"tokens_out":7743,"duration_ms":73609,"significance":"The paper has clear value as a structured inventory: the PRISMA flow, keyword combinations, and architecture-level tables (Tables 3-6) make the corpus and model design space transparent, and the authors are candid in Section 6 about hallucination risk, lack of action validation, and the need for human-in-the-loop supervision. If the corpus-level findings hold, the review provides a useful map of where transformer architectures have matured and where LLM work is still nascent. The Agentic DT vision is a plausible research agenda, and the paper's explicit list of open challenges is a constructive contribution. These strengths are real; however, the empirical generalizations and the concluding 'central insight' need to be rebalanced to match the selective corpus and the vision status of the framework.","major_comments":[{"comment":"In Table 1, the citation-count cutoffs (>=10 for 2021-2022, >=5 for 2023, >=3 for 2024, none for 2025) appear under 'Exclusion Criteria' although the text applies them as minimum inclusion thresholds. Regardless of labeling, these cutoffs are arbitrary and post hoc, and they systematically remove recent, low-cited papers from the corpus. Because Section 7 states as finding (a) that 'Transformer models now dominate energy forecasting tasks,' the review should either rerun the PRISMA selection without the citation filter, provide a sensitivity analysis across thresholds, or explicitly restrict the claim to the 99 selected papers. Without this, the dominance claim may be an artifact of the selection rule.","section":"Section 3, Table 1 and Section 7"},{"comment":"The included corpus is not reconciled with the narrative. Reference [103], a 2021 Prophet-based load forecasting paper, is listed under 'Large Language Model energy prediction' but is neither an LLM nor a transformer study and is never discussed; several other LLM-row entries ([104], [110]-[112]) are also absent from the Section 5 synthesis. This undermines the reproducibility of the screening step and inflates the LLM application counts. The authors should either discuss each included item or document and justify its removal at the synthesis stage.","section":"Table 2, LLM rows, and Section 5"},{"comment":"The paper's central claim about Agentic DTs is presented as a conclusion derived from the review, but the reviewed systems are offline, human-in-the-loop assistants (BEM, retrofit recommendation, forecasting, intrusion detection). None of the 99 studies demonstrates an LLM-driven digital twin executing real-time grid actions or multi-twin negotiation. Moreover, Section 6 itself states that 'LLMs cannot execute or enforce actions directly in the smart grid' and 'humans will remain a key in the decision-making loop,' which directly undercuts the attributes 'autonomous' and 'proactive' attributed to Agentic DTs in Section 7. The conclusion should separate the empirical synthesis (LLMs are at an early, human-supervised stage, concentrated in building energy modeling) from the normative vision (Agentic DT as a research direction with the listed open challenges).","section":"Sections 6-7"}],"minor_comments":[{"comment":"'PRISMA 2000' should be 'PRISMA 2020' in both places; the text cites the 2020 statement, so the label is a typo.","section":"Sections 3 and Figure 3"},{"comment":"The citation-count thresholds should be moved to the inclusion criteria or explicitly labeled as minimum-citation filters; the current placement under 'Exclusion Criteria' makes the procedure ambiguous.","section":"Table 1"},{"comment":"The row for [73] describes an 'LSTM-based Transformer,' while the text in Section 4.1 describes [73] as a 'deep Transformer seq2seq model'; align the terminology.","section":"Section 4.1, Table 3"},{"comment":"Reference formatting is inconsistent: several entries use 'htps://' (missing 's') and some lack complete DOIs (e.g., references [19], [22], [26]); please normalize.","section":"References"},{"comment":"The text says 'more than 20 authors from Europe, more than 10 from the US' and 'dominated by authors from Asia with more than 70 authors,' but the figure appears to count papers or affiliated authors; clarify the unit of analysis.","section":"Figure 7"},{"comment":"The distinction between 'Agentic' and 'Non-Agentic' rows is not defined in the text; add a sentence describing the criterion.","section":"Section 5, Table 7"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is likely to be acceptable in a domain-oriented venue after revision. The review claims are fixable through sensitivity analysis and explicit corpus qualification; the Agentic DT contribution only needs a clear 'vision vs. finding' labeling. I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a competent systematic review with a clear PRISMA search, and the Agentic Digital Twin section is openly a research vision, not a claimed result. The paper does a solid job of mapping transformer applications across load, renewable, and thermal forecasting, and it collects the early LLM uses in buildings and grid management into a readable synthesis. The tables are genuinely helpful for anyone entering the area. The novelty is the framing of LLM-augmented digital twins as agentic—autonomous, proactive, communicative—which is a plausible research agenda, and the challenges list (hallucination, action validation, human-in-the-loop) is honest rather than promotional.\n\nThe main soft spot is methodological. The citation cutoffs by year (10 for 2021–22, 5 for 2023, 3 for 2024, none for 2025) are ad hoc and applied post hoc; they systematically exclude recent low-citation but potentially important work, so the conclusion that transformers now dominate forecasting is at least partly an artifact of the filter. The authors should justify those cutoffs or re-run the synthesis without them. Also, reference [103] is a Prophet-based paper, not an LLM paper, and is never discussed in the text—that looks like a screening error and should be fixed.\n\nThe stress-test concern about extrapolation is fair but somewhat blunted by the paper itself. The leap from offline building energy modeling assistants to autonomous, multi-twin grid management is large, and no reviewed study demonstrates an LLM-driven digital twin operating in a real-time grid. But the authors repeatedly say this is their vision, they note that LLMs cannot enforce actions, and they state that humans remain in the decision loop. So the central claim is over-ambitious relative to the evidence, but it is not presented as a synthesized finding. That is a scope issue, not a hidden circularity or a deceptive result.\n\nThe citation pattern is fine; self-citations are relevant and not excessive. The review ships no code or data, but that is normal for a survey.\n\nWho is this for? Someone looking for an entry point into transformer/LLM energy research, or a reading group discussing how systematic review screening choices shape conclusions. It deserves a serious referee: the review is useful enough to warrant revision, but the authors should fix the corpus issue, justify or remove the citation cutoffs, and explicitly frame the Agentic DT as a research agenda in the abstract and conclusions, not as a near-term development.","headline":"A useful but methodologically filtered review of transformer/LLM energy applications, with the Agentic Digital Twin explicitly a vision rather than a demonstrated result.","tokens_in":37061,"tokens_out":1597,"would_cite":true,"duration_ms":18775,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review of 99 papers argues that transformers have become the standard for energy forecasting and that embedding large language models throughout digital twins will turn them into Agentic Digital Twins—autonomous, proactive…","keywords":["Artificial Intelligence","Transformers","Large Language Models","Generative AI","Agentic Digital Twins","Smart Grid","Energy Forecasting","Building Energy Modeling"],"falsifier":"Re-run the same Web of Science search without the citation-count thresholds and inspect the papers the original rule excluded: if many low-citation 2024–2025 papers show LLMs already operating in grid-management roles beyond building modelling, the review's 'LLMs are still early' conclusion is contradicted; separately, a live microgrid test in which an LLM-augmented digital twin must produce constraint-respecting, actionable recommendations on demand would test whether the Agentic Digital Twin vision has operational substance.","tokens_in":35918,"feed_emoji":"⚡","tokens_out":8104,"duration_ms":71287,"temperature":0.7,"pith_summary":"Drawing on 99 papers published between 2021 and 2025, this systematic review argues that transformer models have become the dominant tool for energy forecasting, consistently outperforming LSTM, GRU, and CNN baselines on demand, thermal load, wind, and solar prediction tasks. It then documents a second, younger wave: large language models are entering energy management not just as forecasters but as integrators of time-series data with text, as generators of synthetic scenarios, and as the engine of multi-agent workflows that automate building-energy modelling and retrofit recommendations. The paper's central, forward-looking claim is that embedding LLMs across every phase of a digital twin will produce 'Agentic Digital Twins'—autonomous, proactive, and communicative models that mirror physical assets, reason about situations, plan energy management actions, and cooperate with other twins. This matters because it would shift AI in smart grids from passive monitoring to active, human-supervised decision support, and it would give grid operators tools for forecasting, balancing, asset onboarding, and workforce training. The authors also identify the obstacles that would have to be overcome for that shift: costly domain fine-tuning, hallucination risk, the need for real-time edge inference, and the absence of standards for LLMs in safety-critical decision loops.","feed_headline":"LLMs turn digital twins into proactive grid agents","feed_subtitle":"A 99-paper review says transformers now rule forecasting while LLMs push twins from passive mirrors to decision-makers","key_machinery":"The load-bearing mechanism is the transformer's self-attention, which computes weighted relationships between all pairs of time steps and lets models capture long-range temporal dependencies; the review shows how domain adaptations—spatial-temporal attention, patch-based decomposition, graph attention, probabilistic decoders, and transfer learning—convert that mechanism into accurate forecasts of electricity demand, thermal load, wind, and solar generation. For the forward-looking half of the paper, the machinery is the fine-tuned LLM inserted into the digital-twin lifecycle: perception, analytics, decision-making, and interaction. The named object that unifies these is the Agentic Digital Twin, an LLM-augmented twin that gains situational awareness, reasoning, proactivity, and social interaction, able to plan and recommend energy actions and negotiate with other twins while remaining subject to human oversight.","core_discovery":"On the paper's own terms, the key discovery is that the energy-AI literature has crossed a threshold. Transformer architectures—above all the Temporal Fusion Transformer, Informer, and patch-based or graph-augmented variants—now define the state of the art in forecasting, while fine-tuned LLMs (GPT-3.5/4, BERT, T5, TimeGPT) and agentic workflows built on them are extending the field beyond prediction into knowledge integration, scenario generation, anomaly detection, and automated simulation. The culminating claim is the Agentic Digital Twin: a next-generation digital twin in which LLMs enhance perception (parsing manuals, reports, and time-series anomalies), analytics (fusing multimodal data with few-shot generalization), decision-making (generating and validating action plans), and interaction (retrieval-augmented knowledge access and cooperation with other twins). In this vision the twin stops being a passive mirror and becomes an active, communicative decision-support agent, with humans remaining in the loop because LLMs cannot enforce actions directly on the grid.","pith_inferences":["If the Agentic Digital Twin vision is right, a likely next bottleneck is multi-twin negotiation: protocols for two LLM-driven twins with conflicting objectives (a building wanting comfort, a grid wanting stability) to reach agreement without a central controller.","The citation-cutoff selection rule may have excluded recent low-citation LLM papers, so the 'LLMs are still early' finding should be re-tested on a search without citation thresholds; the field may be further along than the 99-paper sample suggests.","The building-energy evidence suggests a concrete benchmark: comparing LLM-generated EnergyPlus models against manually authored ones on a standardized test set would quantify how much autonomy agentic twins can actually take on.","A useful refinement of the review's taxonomy would separate 'LLM as forecaster' from 'LLM as orchestrator'—the orchestrator role (agentic workflows, RAG, code generation) is where the promised transformation to Agentic DTs actually lives."],"forward_implications":["Transformer-based forecasting will keep displacing recurrent and statistical baselines in grid operations, making architectures like TFT and Informer the reference points for load, wind, and solar prediction.","LLM-augmented digital twins will likely mature first in building-energy modelling, where multi-agent workflows already automate simulation-file generation, debugging, and retrofit recommendation.","LLMs in grid management will stay advisory rather than directly actuating: hallucination risk and lack of enforcement means humans or dedicated devices execute the recommended actions.","Deploying these models at substations or the edge will require compression and distillation to satisfy real-time latency and privacy constraints.","Standardization, including validation of LLM outputs against simulation engines, will be a precondition for using them in safety-critical decision loops."],"supporting_citations":[{"why":"Supplies the systematic-review method and flow diagram the authors follow to select and screen the 99 papers.","marker":"[34]"},{"why":"Introduces the vanilla transformer and self-attention, the architectural foundation all reviewed forecasting models build on.","marker":"[10]"},{"why":"STELLM shows a pre-trained LLM fine-tuned with spatial-temporal prompts for wind-speed forecasting, evidence that LLMs transfer to energy time series.","marker":"[28]"},{"why":"BERT4ST is an early, high-performing adaptation of BERT to wind-power forecasting, supporting the claim that language models can be specialized for energy tasks.","marker":"[29]"},{"why":"Multi-agent LLM framework that plans and recommends building retrofits, the paper's primary example of agentic LLM workflows.","marker":"[33]"},{"why":"GPT-UBEM demonstrates GPT-4o handling urban building-energy modelling across thousands of buildings, evidence for LLM multimodal data synthesis.","marker":"[102]"},{"why":"Autonomous building-energy-model development and debugging via an LLM agentic workflow, load-bearing for the Agentic Digital Twin's automation claim.","marker":"[106]"},{"why":"Shows LLMs assisting EnergyPlus simulation tasks from input generation to error detection, directly supporting the DT-phase augmentation claim.","marker":"[108]"},{"why":"EPlus-LLM is a fine-tuned T5 that converts natural-language building descriptions into EnergyPlus models, evidence of practical domain adaptation.","marker":"[123]"},{"why":"Frames LLM-enhanced digital-twin modelling, the immediate precedent for the Agentic Digital Twin concept.","marker":"[135]"}],"fun_headline_variants":["LLMs give digital twins a proactive mind for grids","From passive mirrors to proactive agents: LLM-powered twins","Transformers forecast, LLMs act: digital twins go agentic","Agentic digital twins: LLMs drive energy decisions","LLMs transform digital twins into grid decision agents"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's conclusions rest on the assumption that the 99 papers selected through citation-count cutoffs (at least 10 citations for 2021–2022, at least 5 for 2023, at least 3 for 2024, and none for 2025) are representative of the transformer and LLM energy literature; if recent low-citation but significant LLM papers were excluded by those cutoffs, the finding that LLM use is early and concentrated in building modelling could be an artifact of the selection rule rather than the true state of the field.","fun_headline_variants_meta":{"raw":{"variants":["LLMs give digital twins a proactive mind for grids","From passive mirrors to proactive agents: LLM-powered twins","Transformers forecast, LLMs act: digital twins go agentic","Agentic digital twins: LLMs drive energy decisions","LLMs transform digital twins into grid decision agents"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000287,"raw_usage":{"total_tokens":1729,"prompt_tokens":1030,"completion_tokens":699,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":620}},"tokens_in":646,"tokens_out":699,"duration_ms":7286,"temperature":1.0,"reasoning_tokens":620,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:17:38.389642+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same Web of Science search without the citation-count thresholds and inspect the papers the original rule excluded: if many low-citation 2024–2025 papers show LLMs already operating in grid-management roles beyond building modelling, the review's 'LLMs are still early' conclusion is contradicted; separately, a live microgrid test in which an LLM-augmented digital twin must produce constraint-respecting, actionable recommendations on demand would test whether the Agentic Digital Twin vision has operational substance.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the systematic-review method and flow diagram the authors follow to select and screen the 99 papers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"STELLM shows a pre-trained LLM fine-tuned with spatial-temporal prompts for wind-speed forecasting, evidence that LLMs transfer to energy time series."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"BERT4ST is an early, high-performing adaptation of BERT to wind-power forecasting, supporting the claim that language models can be specialized for energy tasks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"GPT-UBEM demonstrates GPT-4o handling urban building-energy modelling across thousands of buildings, evidence for LLM multimodal data synthesis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Autonomous building-energy-model development and debugging via an LLM agentic workflow, load-bearing for the Agentic Digital Twin's automation claim."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows LLMs assisting EnergyPlus simulation tasks from input generation to error detection, directly supporting the DT-phase augmentation claim."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"EPlus-LLM is a fine-tuned T5 that converts natural-language building descriptions into EnergyPlus models, evidence of practical domain adaptation."}],"review_version":1}