{"id":"53e81497-60d1-4d59-a66b-dc0b967b56fb","arxiv_id":"2607.03176","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"AIRL-recovered reward functions show that the 2022 energy crisis and heatwave reshaped Italian households' cooling responses heterogeneously, sometimes durably, with time-of-use as an independent dimension.","lead":"Italian household cooling electricity use is cast as agent behaviour whose reward functions are recovered by Adversarial Inverse Reinforcement Learning from smart-meter data. The 2022 energy-price shock and heatwave reshaped those rewards differently across consumer groups, with some shifts lasting into 2023.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The single-behaviour-per-cluster assumption is load-bearing for the heterogeneous-response claim, and the paper never checks whether recovered rewards are stable under re-clustering or multi-intention alternatives.","rationale":"The Reader correctly flags the single-behaviour-per-cluster assumption (§4.2.1) as the weakest link. That assumption is not a peripheral modelling convenience: the entire interpretive apparatus of §5.2–5.3 (persistent vs transient vs negligible reward shifts, and the claim that time-of-use is an independent dimension) rests on each recovered f_φ being the reward of a coherent agent class. The paper supplies no robustness check against this assumption, even though its own sub-clustering already demonstrates residual behavioural heterogeneity inside the high-consumption group. Private data and the large hyper-parameter surface further prevent external verification. These points keep the verdict at CONDITIONAL rather than ACCEPT; they do not justify REJECT, because the model-free empirical consumption curves move in the same qualitative directions as the rewards, and the engineering contribution remains real. The concrete test above would settle whether the spectrum is robust or partition-dependent; until it is run, the central claim stays conditional on an untested modelling premise.","tokens_in":23957,"tokens_out":718,"duration_ms":7103,"concrete_test":"Re-run the hierarchical clustering of §4.1 with three alternative variance splits (30/70, 50/50, 70/30 static vs dynamic) and with pure consumption-profile k-means; retrain AIRL on the resulting high-consumption and urban-intermediate partitions for 2021–2023. If the qualitative labels (persistent decline, transient amplification, negligible change) flip for any of the four focal groups, or if the afternoon-high vs evening-high divergence of Fig. 10 disappears under the original weighting, the spectrum claim is partition-dependent and the single-behaviour assumption fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim (Abstract, §5.3) is that shocks reshaped cooling behaviour heterogeneously across groups, producing a spectrum of transient / durable / negligible responses, with time-of-use as a separate dimension. That claim is only interpretable if each cluster is a single coherent behaviour, as the authors state in §4.2.1: “By clustering on all variables at once, we also assume that each cluster has an internally coherent environment and represents a single behaviour … avoiding the need of multiple-intention modeling.” If a cluster mixes latent intentions (e.g., AC owners vs non-owners, or day-active vs evening-active households that the hierarchical clustering failed to separate), the AIRL reward is an average that need not correspond to any real household policy. The sub-clustering exercise in Fig. 10 already shows that, once variance weight is shifted toward consumption timing, two high-consumption groups that look similar on context variables diverge sharply in 2022 response; this is evidence that the original 50/50 static/dynamic weighting can leave residual multi-intention structure inside a cluster. Because the paper never reports reward stability under alternative clusterings, different variance weights, or a multi-intention IRL baseline, the spectrum of responses could be an artefact of the clustering partition rather than a property of real households.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes an Adversarial Inverse Reinforcement Learning (AIRL) pipeline to recover model-implied reward functions for household electricity consumption from Italian smart-meter data (May–July 2021–2023). Households are clustered via Echo State Networks, Tensor PCA and hierarchical clustering on static socioeconomic/environmental attributes and dynamic UTCI–consumption series; each selected cluster is treated as a single-behaviour expert and trained with AIRL (PPO generator, hinge-loss discriminator with spectral/batch norm). Reward surfaces and empirical consumption-vs-UTCI curves are then compared across pre-crisis, crisis and post-peak summers for four clusters (high-consumption, urban intermediate, rural low, urban low) and two high-consumption sub-clusters differing in daily timing. The authors report a spectrum of cooling-behaviour responses—persistent, transient and negligible—conditioned by prior habits, built environment and time-of-use, and argue that demand-response design should account for who, where, when and persistence of shock response.","tokens_in":24346,"tokens_out":1636,"duration_ms":20154,"significance":"If the recovered rewards are reliable behavioural summaries, the work offers a transferable, non-linear alternative to setpoint/slope and survey-based characterisations of cooling under concurrent price and heat stress, with direct policy relevance for demand response. Strengths include: (i) joint use of reward surfaces and model-free empirical consumption–UTCI curves; (ii) OOD validation by freezing the reward and re-fitting the policy on adjacent April/August windows; (iii) an explicit sub-daily time-of-use analysis that isolates timing as a dimension orthogonal to socioeconomic context; and (iv) a carefully engineered AIRL stack (hinge loss, spectral norm, action-scale annealing, noise injection). These elements go beyond typical black-box load forecasting and beyond most energy-poverty cooling studies that lack quantified demand responses under the 2022 crisis.","major_comments":[{"comment":"§4.2.1 states that clustering on all variables at once implies each cluster has an ‘internally coherent environment and represents a single behaviour,’ thereby avoiding multi-intention IRL. This assumption is load-bearing for interpreting f_ϕ as a household-level cooling policy. Fig. 10 already shows that reweighting variance toward consumption timing splits the high-consumption cluster into afternoon-high vs evening-high groups with sharply different 2022 responses, indicating residual multi-intention structure under the original 50/50 split. The manuscript never reports reward stability under alternative variance weights, different numbers of clusters, or a multi-intention IRL baseline. Without such checks, the claimed spectrum (persistent / transient / negligible) could partly be an artefact of the partition rather than a property of real households. A minimal robustness suite—re-clus","section":null},{"comment":"§5.1 obtains 11 clusters but applies AIRL only to four (plus two sub-clusters), selected ‘to highlight a good share of the behavioural differences’ with a cooling focus. The heterogeneous-response claim is therefore conditioned on a non-random subset. Either (a) report reward surfaces for the remaining clusters (or a random sample of them) to show the spectrum is not selection-driven, or (b) pre-specify selection criteria (e.g., AC ownership quantiles, urban/rural extremes) and justify why intermediate clusters 3–6, 8, 10–11 would not alter the typology. As written, external validity of the three-way typology is unclear.","section":null},{"comment":"§4.3 calibrates a softmax temperature τ so that the expected consumption under the reward matches the empirical mean, then visualises E_τ[e|s] vs UTCI. Year-to-year comparisons of these surfaces (Figs. 6–9) are purely qualitative—no distance metric, confidence bands on the reward surface, or formal test of whether 2022 vs 2021 (or 2023 vs 2021) curves differ. Given partial identification of rewards (§4.2.1, §5.4), absolute levels are not meaningful; only comparative statements are. The paper should define a quantitative comparison (e.g., integrated absolute difference of calibrated surfaces above 30° UTCI, or a bootstrap over Monte-Carlo state samples) and report it for each cluster/year pair that underpins the persistent/transient/negligible labels.","section":null},{"comment":"Several free parameters that shape both clustering and reward recovery are fixed without sensitivity analysis: the 50/50 static–dynamic variance split (§4.1), Combined Metric weights α,β,γ,η,ζ (§4.2.2), action-scale annealing and discriminator noise scale (§4.2.3), and the number of retained clusters/sub-clusters. Because the central claim is about heterogeneous behavioural change, at least the variance split and CM weights should be varied and the qualitative typology re-checked. If the typology is stable, that strengthens the result; if not, the free-parameter dependence must be disclosed as a limitation of the spectrum claim.","section":null}],"minor_comments":[{"comment":"§3: AC ownership is estimated on 2021 data and held fixed through 2022–2023 to avoid crisis-period bias. This is reasonable but should be flagged more prominently in §5.3 when interpreting high-AC vs low-AC clusters, since true ownership may have changed.","section":null},{"comment":"Figure 5 axis labels and cluster ordering are hard to parse in the text rendering; ensure the published figures have legible tick labels and a clear legend for the four selected clusters.","section":null},{"comment":"§4.2.1: the reward decomposition r(s,a,s′)=f_ϕ(s,a)+γΦ_ψ(s′)−Φ_ψ(s) is standard AIRL but is not numbered as an equation; numbering it would help cross-reference in §5.4’s partial-identification discussion.","section":null},{"comment":"Related Work §2 cites Fu et al. (2022) under the AIRL reference [14]; that paper is a review of RL for building control, not the original AIRL paper (Fu, Luo & Levine, 2018, arXiv:1710.11248, already listed as [48]). Correct the citation mapping.","section":null},{"comment":"§5.4 mentions LSTM generators and environment discretisation as future work; a short note on current CM values (or a table in the supplement) for the four main clusters would help readers judge how well the present continuous PPO generator already reproduces expert trajectories.","section":null},{"comment":"Abstract and §1 use both ‘sub-daily’ and ‘intradaily’; pick one term for consistency.","section":null}],"recommendation":"major_revision","confidential_remarks":"The application of AIRL to large-scale smart-meter behavioural analysis is novel for this journal’s audience and the empirical side-by-side design is a genuine strength. The main risk is over-interpretation of partially identified rewards under an untested single-behaviour clustering assumption. If the authors deliver the robustness checks in major comments 1–4, the paper would be a solid contribution; if they cannot, the claims should be narrowed to the model-free empirical curves plus a more cautious IRL illustration. Scope fit for a methods/energy-systems venue is good; pure ML venues may want stronger multi-intention baselines."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is that this is a careful, large-scale application of AIRL to residential smart-meter data that recovers cooling reward surfaces before, during and after the 2022 price-and-heat shock, and that the recovered surfaces line up with the model-free empirical consumption-vs-UTCI curves. That combination is new relative to the survey, set-point and small-sample IRL literature they cite.\n\nWhat they do well: the pipeline is engineered with care (ESN+Tensor-PCA clustering, hinge-loss AIRL with spectral norm and OOD policy re-fit under frozen reward). They report both the reward surfaces and the raw empirical curves side-by-side, so the qualitative spectrum—persistent, transient, negligible—is not an artefact of the reward alone. The sub-clustering on timing (afternoon-high vs evening-high) is a clean demonstration that time-of-use is a separate behavioural axis even when socioeconomic and environmental covariates look similar. Citations are appropriate; the math is standard AIRL with sensible stabilisers.\n\nThe soft spot is real but not fatal. They explicitly assume each cluster is a single coherent behaviour (§4.2.1). Their own Fig. 10 shows that shifting the variance weight toward consumption timing splits the high-consumption group into two groups that respond differently in 2022. That is evidence residual multi-intention structure can remain. They never report reward stability under re-clustering, different variance splits, or a multi-intention baseline. So the “spectrum of responses” could be partly partition-dependent. Private data and a large hyper-parameter surface are secondary practical limits, not conceptual ones.\n\nThis is for energy-systems and climate-adaptation people who need transferable behavioural response surfaces, and for IRL practitioners looking for a non-robotics application. It deserves a serious referee. I would engage with it once code and anonymised data appear; the central qualitative finding is already usable.","headline":"Solid new application of AIRL to Italian smart-meter cooling behaviour across the energy crisis; the heterogeneous-response claim is real but rests on a single-behaviour-per-cluster assumption that is only partially stress-tested.","tokens_in":24939,"tokens_out":493,"would_cite":true,"duration_ms":5700,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Energy crisis and heat reshaped Italian cooling behaviour differently by group: some shifts stuck, some faded, some never appeared.","keywords":["Inverse Reinforcement Learning","household electricity demand","cooling behaviour","energy crisis","smart meter data","time-of-use","AIRL","thermal stress"],"falsifier":"Re-run the same AIRL pipeline after deliberately splitting a high-consumption cluster by latent intention or by a held-out socioeconomic split; if the recovered 2021–2023 reward trajectories reverse or collapse into noise, the single-behaviour-per-cluster claim fails.","tokens_in":24858,"feed_emoji":"⚡","tokens_out":904,"duration_ms":8675,"temperature":0.7,"pith_summary":"This paper treats Italian households as agents and recovers their electricity-use behaviour as reward functions via Adversarial Inverse Reinforcement Learning on smart-meter data. The aim is to show how cooling-related consumption responds to temperature when households face concurrent socioeconomic and climatic shocks. Using May–July smart-meter data for 2021–2023—before, during, and after the European energy-price spike and a severe heatwave—the authors cluster consumers by socioeconomic context, built environment, and load profiles, then recover each cluster’s reward. The recovered rewards reveal a spectrum of responses: high-consumption users cut high-temperature use and kept much of that cut into 2023; urban intermediate users amplified then largely reversed; rural low-consumption users gradually strengthened temperature response; urban low-consumption users barely moved. Groups that look similar on income and location but differ in daily timing of use also diverge, so time-of-use is treated as its own behavioural dimension. The practical claim is that demand-response and energy policy must track who people are, where they live, when they consume, and whether a shock-induced change lasts.","feed_headline":"Crisis and heat rewired cooling habits—some cuts stuck","feed_subtitle":"Italian smart-meter rewards show durable, transient, and null shifts by group and daily timing","key_machinery":"Adversarial Inverse Reinforcement Learning (AIRL) reward surfaces: each cluster is treated as an agent whose actions are consumption variations, and AIRL recovers a model-implied reward function that is then sampled into optimal consumption-versus-UTCI curves comparable across years and clusters.","core_discovery":"Socioeconomic and climatic shocks of 2021–2023 reshaped cooling behaviour heterogeneously across Italian consumer clusters, in directions set by prior habits and built environment, producing durable, transient, and negligible shifts; within high-consumption users, groups that differ only in daily timing of use also respond differently, so time-of-use is a separate axis of heterogeneity.","pith_inferences":["If durable high-consumption cuts reflect permanent habit change rather than temporary thrift, post-crisis rebound forecasts that assume 2021 elasticities will overstate summer peak load.","Sub-daily timing differences that survive socioeconomic matching suggest activity schedules are a policy lever for cooling demand response comparable to price signals.","Partial identification of rewards means absolute reward levels should not be used for welfare ranking; only cross-year and cross-cluster shape comparisons are licensed by the method.","Extending the same AIRL pipeline to heating seasons or to regions with different AC penetration would test whether the durable/transient/negligible spectrum generalises beyond Italian summers."],"forward_implications":["Demand-response and tax schemes must condition on whether a group’s response to a shock is durable or reverts once prices and temperatures ease.","Policies that ignore daily timing of use will mis-target high-consumption households that look socioeconomically similar.","Low-consumption rural groups can grow cooling response under heat even while urban low-consumption peers stay flat, so equity and grid planning need location-specific trajectories.","Year-to-year temperature-response functions cannot be treated as fixed parameters in long-term energy and climate models.","Smart-meter analysis that recovers rewards can separate temperature response from time-of-day confounding that raw averages hide."],"fun_headline_variants":["Shocks rewired cooling habits into durable transient or null shifts","Crisis and heat left lasting fleeting or zero cooling changes by group","Prior habits set which cooling rewirings stuck after 2021-23 shocks","Daily timing split cooling responses among similar Italian clusters","IRL maps uneven cooling reward shifts under crisis and heatwave"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Clustering on all variables at once is taken to give each group one coherent environment and one behaviour, so a single recovered reward can stand for that group.","fun_headline_variants_meta":{"raw":{"variants":["Shocks rewired cooling habits into durable transient or null shifts","Crisis and heat left lasting fleeting or zero cooling changes by group","Prior habits set which cooling rewirings stuck after 2021-23 shocks","Daily timing split cooling responses among similar Italian clusters","IRL maps uneven cooling reward shifts under crisis and heatwave"]},"model":"grok-4.5","effort":"low","cost_usd":0.009816,"raw_usage":{"total_tokens":2244,"prompt_tokens":816,"num_sources_used":0,"completion_tokens":88,"cost_in_usd_ticks":98160000,"prompt_tokens_details":{"text_tokens":816,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1340,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":816,"tokens_out":88,"duration_ms":10801,"temperature":1.0,"reasoning_tokens":1340,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T04:19:50.688085+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same AIRL pipeline after deliberately splitting a high-consumption cluster by latent intention or by a held-out socioeconomic split; if the recovered 2021–2023 reward trajectories reverse or collapse into noise, the single-behaviour-per-cluster claim fails.","supporting_citations":[],"review_version":1}