REVIEW 4 major objections 6 minor 67 references
Energy crisis and heat reshaped Italian cooling behaviour differently by group: some shifts stuck, some faded, some never appeared.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 04:19 UTC pith:6KYAANOS
load-bearing objection Solid new application of AIRL to Italian smart-meter cooling behaviour across the energy crisis; the heterogeneous-response claim is real but rests on a single-behaviour-per-cluster assumption that is only partially stress-tested. the 4 major comments →
Understanding electricity consumption behaviour through Inverse Reinforcement Learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Socioeconomic and climatic shocks of 2021–2023 reshaped cooling behaviour heterogeneously across Italian consumer clusters, in directions set by prior habits and built environment, producing durable, transient, and negligible shifts; within high-consumption users, groups that differ only in daily timing of use also respond differently, so time-of-use is a separate axis of heterogeneity.
What carries the argument
Adversarial Inverse Reinforcement Learning (AIRL) reward surfaces: each cluster is treated as an agent whose actions are consumption variations, and AIRL recovers a model-implied reward function that is then sampled into optimal consumption-versus-UTCI curves comparable across years and clusters.
Load-bearing premise
Clustering on all variables at once is taken to give each group one coherent environment and one behaviour, so a single recovered reward can stand for that group.
What would settle it
Re-run the same AIRL pipeline after deliberately splitting a high-consumption cluster by latent intention or by a held-out socioeconomic split; if the recovered 2021–2023 reward trajectories reverse or collapse into noise, the single-behaviour-per-cluster claim fails.
If this is right
- Demand-response and tax schemes must condition on whether a group’s response to a shock is durable or reverts once prices and temperatures ease.
- Policies that ignore daily timing of use will mis-target high-consumption households that look socioeconomically similar.
- Low-consumption rural groups can grow cooling response under heat even while urban low-consumption peers stay flat, so equity and grid planning need location-specific trajectories.
- Year-to-year temperature-response functions cannot be treated as fixed parameters in long-term energy and climate models.
- Smart-meter analysis that recovers rewards can separate temperature response from time-of-day confounding that raw averages hide.
Where Pith is reading between the lines
- If durable high-consumption cuts reflect permanent habit change rather than temporary thrift, post-crisis rebound forecasts that assume 2021 elasticities will overstate summer peak load.
- Sub-daily timing differences that survive socioeconomic matching suggest activity schedules are a policy lever for cooling demand response comparable to price signals.
- Partial identification of rewards means absolute reward levels should not be used for welfare ranking; only cross-year and cross-cluster shape comparisons are licensed by the method.
- Extending the same AIRL pipeline to heating seasons or to regions with different AC penetration would test whether the durable/transient/negligible spectrum generalises beyond Italian summers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an Adversarial Inverse Reinforcement Learning (AIRL) pipeline to recover model-implied reward functions for household electricity consumption from Italian smart-meter data (May–July 2021–2023). Households are clustered via Echo State Networks, Tensor PCA and hierarchical clustering on static socioeconomic/environmental attributes and dynamic UTCI–consumption series; each selected cluster is treated as a single-behaviour expert and trained with AIRL (PPO generator, hinge-loss discriminator with spectral/batch norm). Reward surfaces and empirical consumption-vs-UTCI curves are then compared across pre-crisis, crisis and post-peak summers for four clusters (high-consumption, urban intermediate, rural low, urban low) and two high-consumption sub-clusters differing in daily timing. The authors report a spectrum of cooling-behaviour responses—persistent, transient and negligible—conditioned by prior habits, built environment and time-of-use, and argue that demand-response design should account for who, where, when and persistence of shock response.
Significance. If the recovered rewards are reliable behavioural summaries, the work offers a transferable, non-linear alternative to setpoint/slope and survey-based characterisations of cooling under concurrent price and heat stress, with direct policy relevance for demand response. Strengths include: (i) joint use of reward surfaces and model-free empirical consumption–UTCI curves; (ii) OOD validation by freezing the reward and re-fitting the policy on adjacent April/August windows; (iii) an explicit sub-daily time-of-use analysis that isolates timing as a dimension orthogonal to socioeconomic context; and (iv) a carefully engineered AIRL stack (hinge loss, spectral norm, action-scale annealing, noise injection). These elements go beyond typical black-box load forecasting and beyond most energy-poverty cooling studies that lack quantified demand responses under the 2022 crisis.
major comments (4)
- §4.2.1 states that clustering on all variables at once implies each cluster has an ‘internally coherent environment and represents a single behaviour,’ thereby avoiding multi-intention IRL. This assumption is load-bearing for interpreting f_ϕ as a household-level cooling policy. Fig. 10 already shows that reweighting variance toward consumption timing splits the high-consumption cluster into afternoon-high vs evening-high groups with sharply different 2022 responses, indicating residual multi-intention structure under the original 50/50 split. The manuscript never reports reward stability under alternative variance weights, different numbers of clusters, or a multi-intention IRL baseline. Without such checks, the claimed spectrum (persistent / transient / negligible) could partly be an artefact of the partition rather than a property of real households. A minimal robustness suite—re-clus
- §5.1 obtains 11 clusters but applies AIRL only to four (plus two sub-clusters), selected ‘to highlight a good share of the behavioural differences’ with a cooling focus. The heterogeneous-response claim is therefore conditioned on a non-random subset. Either (a) report reward surfaces for the remaining clusters (or a random sample of them) to show the spectrum is not selection-driven, or (b) pre-specify selection criteria (e.g., AC ownership quantiles, urban/rural extremes) and justify why intermediate clusters 3–6, 8, 10–11 would not alter the typology. As written, external validity of the three-way typology is unclear.
- §4.3 calibrates a softmax temperature τ so that the expected consumption under the reward matches the empirical mean, then visualises E_τ[e|s] vs UTCI. Year-to-year comparisons of these surfaces (Figs. 6–9) are purely qualitative—no distance metric, confidence bands on the reward surface, or formal test of whether 2022 vs 2021 (or 2023 vs 2021) curves differ. Given partial identification of rewards (§4.2.1, §5.4), absolute levels are not meaningful; only comparative statements are. The paper should define a quantitative comparison (e.g., integrated absolute difference of calibrated surfaces above 30° UTCI, or a bootstrap over Monte-Carlo state samples) and report it for each cluster/year pair that underpins the persistent/transient/negligible labels.
- Several free parameters that shape both clustering and reward recovery are fixed without sensitivity analysis: the 50/50 static–dynamic variance split (§4.1), Combined Metric weights α,β,γ,η,ζ (§4.2.2), action-scale annealing and discriminator noise scale (§4.2.3), and the number of retained clusters/sub-clusters. Because the central claim is about heterogeneous behavioural change, at least the variance split and CM weights should be varied and the qualitative typology re-checked. If the typology is stable, that strengthens the result; if not, the free-parameter dependence must be disclosed as a limitation of the spectrum claim.
minor comments (6)
- §3: AC ownership is estimated on 2021 data and held fixed through 2022–2023 to avoid crisis-period bias. This is reasonable but should be flagged more prominently in §5.3 when interpreting high-AC vs low-AC clusters, since true ownership may have changed.
- Figure 5 axis labels and cluster ordering are hard to parse in the text rendering; ensure the published figures have legible tick labels and a clear legend for the four selected clusters.
- §4.2.1: the reward decomposition r(s,a,s′)=f_ϕ(s,a)+γΦ_ψ(s′)−Φ_ψ(s) is standard AIRL but is not numbered as an equation; numbering it would help cross-reference in §5.4’s partial-identification discussion.
- Related Work §2 cites Fu et al. (2022) under the AIRL reference [14]; that paper is a review of RL for building control, not the original AIRL paper (Fu, Luo & Levine, 2018, arXiv:1710.11248, already listed as [48]). Correct the citation mapping.
- §5.4 mentions LSTM generators and environment discretisation as future work; a short note on current CM values (or a table in the supplement) for the four main clusters would help readers judge how well the present continuous PPO generator already reproduces expert trajectories.
- Abstract and §1 use both ‘sub-daily’ and ‘intradaily’; pick one term for consistency.
Circularity Check
No load-bearing circular derivation: AIRL rewards are descriptive recoveries from trajectories, with independent empirical curves and OOD checks; only a minor mean-calibration of the softmax temperature is by construction.
specific steps
-
fitted input called prediction
[§4.3 Behaviour analysis (softmax calibration of τ)]
"a softmax distribution over consumption values is constructed at temperature τ: π_τ(e|s)∝exp(R_θ(s[e],0)/τ) ... and τ is found so that 1/N ∑_{i=1}^N E_τ[e|s_i] = ē_emp, with ē_emp the empirical mean consumption. The calibrated surface E_τ[e|s] constitutes the primary output for the behaviour analysis."
The scalar τ is fitted so the average of the reward-implied consumption equals the empirical mean by construction. Absolute level of the ‘optimal consumption’ surface is therefore not an independent prediction. This is minor: one scalar cannot force the shape of the response versus UTCI that underpins the transient/durable/negligible claims, and empirical curves are reported separately without AIRL.
full rationale
The paper’s central claims (heterogeneous transient/durable/negligible cooling-response shifts across clusters and years; time-of-use as a separate dimension) are not forced by definition or by a self-citation uniqueness chain. AIRL recovers a model-implied reward from observed trajectories—that is the method, not a first-principles prediction that reduces to its inputs. The authors explicitly treat the recovered f_ϕ as only partially identified and not a psychological preference measure, report model-free empirical consumption-vs-UTCI curves side-by-side with the reward surfaces, and freeze the reward for OOD policy re-fit on adjacent weeks. There is no uniqueness theorem imported from overlapping authors that forbids alternatives, and no renaming of a known empirical pattern as a derived law. The single-behaviour-per-cluster assumption (§4.2.1) is load-bearing for interpretation and is a validity risk (as the sub-clustering already hints), but it is an assumption about partition coherence, not a circular reduction of the claimed spectrum to the clustering inputs. The only mild by-construction step is the scalar temperature τ chosen so that the softmax-implied mean consumption matches the empirical mean; that fixes the overall level, not the shape of the UTCI response that carries the heterogeneity claims. Overall circularity is therefore negligible.
Axiom & Free-Parameter Ledger
free parameters (5)
- clustering variance split (static vs dynamic)
- softmax temperature τ
- AIRL Combined Metric weights (α,β,γ,η,ζ)
- action-scale annealing schedule and discriminator noise scale
- number of clusters / sub-clusters retained for analysis
axioms (4)
- domain assumption Observed trajectories are generated by agents that act (near-)optimally with respect to an unknown reward in a Markov Decision Process.
- ad hoc to paper Each cluster possesses an internally coherent environment and represents a single latent behaviour, so multi-intention IRL is unnecessary.
- domain assumption AC ownership estimated on 2021 data remains constant through 2022–2023.
- standard math The potential-based shaping term can be discarded at evaluation time, leaving only f_ϕ as the transferable reward.
invented entities (1)
-
model-implied reward function f_ϕ as compact representation of household cooling behaviour
no independent evidence
read the original abstract
Understanding how households consume electricity in response to socioeconomic and climatic drivers is important for decision-makers designing energy policies in a changing climate and under geopolitical tensions. Consumers respond differently to thermal stress depending on income, consumption habits and the surrounding built environment, a nonlinear behaviour that most approaches oversimplify. In this study, households are treated as agents interacting with complex environments, and Inverse Reinforcement Learning is used to represent their consumption behaviour as model implied reward functions. Specifically, we observe how these reward functions change when households undergo socioeconomic and climatic shocks. The framework is tested on different clusters of electricity consumption profiles in Italy. Clusters' reward functions are retrieved and used to understand how cooling behaviour changes from summer 2021 to summer 2022 and 2023, before, during and after the energy crisis and a heatwave. We find that these shocks reshaped cooling behaviour heterogeneously across consumer groups, in directions conditioned by their prior habits and built environment. Across the 2021 to 2023 summers, we identify a spectrum of responses: transient adjustments that receded as the shocks eased, durable shifts persisting into 2023, and consumers exhibiting negligible change. At the intradaily scale, groups comparable in socioeconomic and environmental context but differing in their daily timing of consumption responded distinctly, identifying time of use as a separate dimension of behavioural heterogeneity. Energy policies and demand-response schemes should therefore account not only for who consumers are and where they live, but for when they consume and whether their response to a shock persists.
Figures
Reference graph
Works this paper leans on
-
[1]
953–1048.doi:10.1017/ 9781009157926.011
Buildings, in: Intergovernmental Panel On Climate Change (Ipcc) (Ed.), Climate Change 2022 - Miti- gation of Climate Change, 1st Edition, Cambridge University Press, 2023, pp. 953–1048.doi:10.1017/ 9781009157926.011
2022
-
[2]
S. Cong, D. Nock, Y. L. Qiu, B. Xing, Unveiling hidden energy poverty using the energy equity gap, Nature Communications 13 (1) (2022) 2456.doi: 10.1038/s41467-022-30146-5. 14
-
[3]
P. Andersen, A. M. Dietrich, Price response in res- idential electricity demand: Evidence from Danish smart meter data, Energy Economics 153 (2026) 109087.doi:10.1016/j.eneco.2025.109087
-
[4]
A. Ushakova, S. Jankin Mikhaylov, Big data to the rescue? Challenges in analysing granular household electricity consumption in the United Kingdom, En- ergy Research & Social Science 64 (2020) 101428. doi:10.1016/j.erss.2020.101428
-
[5]
S. Vitiello, N. Andreadou, M. Ardelean, G. Fulli, Smart Metering Roll-Out in Europe: Where Do We Stand? Cost Benefit Analyses in the Clean Energy Package and Research Trends in the Green Deal, En- ergies 15 (7) (2022) 2340.doi:10.3390/en15072340
-
[6]
Happle, J
G. Happle, J. A. Fonseca, A. Schlueter, A review on occupant behavior in urban building energy models, Energy and Buildings 174 (2018) 276–292.doi:10. 1016/j.enbuild.2018.06.030
2018
-
[7]
G. M. Huebner, M. McMichael, D. Shipworth, M. Shipworth, M. Durand-Daubin, A. J. Summer- field, The shape of warmth: Temperature profiles in living rooms, Building Research & Information 43 (2) (2015) 185–196.doi:10.1080/09613218. 2014.922339
-
[8]
Oreane.Y. Edelenbosch, L. Miu, J. Sachs, A. Hawkes, M.Tavoni, Translatingobservedhouseholdenergybe- havior to agent-based technology choices in an inte- grated modeling framework, iScience 25 (3) (2022) 103905.doi:10.1016/j.isci.2022.103905
-
[9]
O. Motlagh, P. Paevere, T. S. Hong, G. Grozev, Anal- ysis of household electricity consumption behaviours: Impact of domestic electricity generation, Applied Mathematics and Computation 270 (2015) 165–178. doi:10.1016/j.amc.2015.08.029
-
[10]
H. Uchida, K. Kishimoto, K. Nishizawa, Y. Shimoda, Y. Yamaguchi, K. Togawa, Aggregated smart me- ter data driven occupant behavior analysis based on inverse problem optimization, Energy and Buildings 345 (2025) 116074.doi:10.1016/j.enbuild.2025. 116074
-
[11]
J. Einolander, A. Kiviaho, R. Lahdelma, Detecting changes in price-sensitivity of household electricity consumption: The impact of the global energy crisis on implicit demand response behavior of Finnish de- tached households, Energy and Buildings 306 (2024) 113941.doi:10.1016/j.enbuild.2024.113941
-
[12]
S. Harish, N. Singh, R. Tongia, Impact of temper- ature on electricity demand: Evidence from Delhi and Indian states, Energy Policy 140 (2020) 111445. doi:10.1016/j.enpol.2020.111445
-
[13]
A. Y. Ng, S. J. Russell, Algorithms for Inverse Re- inforcement Learning, in: Proceedings of the Seven- teenth International Conference on Machine Learn- ing, ICML ’00, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 2000, pp. 663–670
2000
-
[14]
Q. Fu, Z. Han, J. Chen, Y. Lu, H. Wu, Y. Wang, Ap- plications of reinforcement learning for building en- ergy efficiency control: A review, Journal of Building Engineering50(2022)104165.doi:10.1016/j.jobe. 2022.104165
-
[15]
M. Sackmann, H. Bey, U. Hofmann, J. Thielecke, Modeling Driver Behavior using Adversarial Inverse Reinforcement Learning, in: 2022 IEEE Intelligent Vehicles Symposium (IV), 2022, pp. 1683–1690.doi: 10.1109/IV51971.2022.9827292
-
[16]
P. Wang, D. Liu, J. Chen, H. Li, C.-Y. Chan, Deci- sion Making for Autonomous Driving via Augmented Adversarial Inverse Reinforcement Learning, in: 2021 IEEE International Conference on Robotics and Au- tomation (ICRA), 2021, pp. 1036–1042.doi:10. 1109/ICRA48506.2021.9560907
arXiv 2021
-
[17]
I. Staffell, S. Pfenninger, N. Johnson, A global model of hourly space heating and cooling demand at multi- ple spatial scales, Nature Energy 8 (12) (2023) 1328– 1344.doi:10.1038/s41560-023-01341-5
-
[19]
J. Burkhardt, K. T. Gillingham, P. K. Kopalle, Field Experimental Evidence on the Effect of Pric- ing on Residential Electricity Conservation, Manage- ment Science 69 (12) (2023) 7784–7798.doi:10. 1287/mnsc.2020.02074
arXiv 2023
-
[20]
Ahlvik, T
L. Ahlvik, T. Kaariaho, M. Liski, I. Vehviläinen, Household-Level Responses to the European Energy Crisis (2025)
2025
-
[21]
J. Lunghi, J. Bonan, C. Cattaneo, G. d’Adda, M. Tavoni, Power Play: Balancing Efficiency and Protection in Fixed vs. Variable Electricity Pricing (Jan. 2026).arXiv:6064532,doi:10.2139/ssrn. 6064532
doi:10.2139/ssrn 2026
-
[22]
Y. Peng, C. A. Klöckner, Drivers and barriers to energy-saving behaviour formation and retention in response to extreme events: Insights from the energy crisis, Sustainability Science 21 (2) (2026) 637–655. doi:10.1007/s11625-025-01764-x. 15
-
[23]
E. De Cian, G. Falchetta, F. Pavanello, Y. Romitti, I. Sue Wing, The impact of air conditioning on res- idential electricity consumption across world coun- tries, Journal of Environmental Economics and Man- agement 131 (2025) 103122.doi:10.1016/j.jeem. 2025.103122
-
[24]
G. Falchetta, E. D. Cian, F. Pavanello, I. S. Wing, Inequalities in global residential cooling energy use to 2050, Nature Communications 15 (1) (2024) 7874. doi:10.1038/s41467-024-52028-8
-
[25]
Z. Wang, B. Lu, B. Wang, Y. L. Qiu, H. Shi, B. Zhang, J. Li, H. Li, W. Zhao, Incentive based emergency demand response effectively reduces peak load during heatwave without harm to vulnerable groups, Nature Communications 14 (1) (2023) 6202. doi:10.1038/s41467-023-41970-8
-
[26]
M. Kwon, S. Cong, D. Nock, L. Huang, Y. L. Qiu, B. Xing, Forgone summertime comfort as a function of avoided electricity use, Energy Policy 183 (2023) 113813.doi:10.1016/j.enpol.2023.113813
-
[27]
L. Huang, D. Nock, S. Cong, Y. L. Qiu, Inequali- ties across cooling and heating in households: En- ergy equity gaps, Energy Policy 182 (2023) 113748. doi:10.1016/j.enpol.2023.113748
-
[28]
E. Chatzikonstantinou, N. Katsoulakos, F. Vatavali, Housing and energy consumption in Greece. House- holds’ experiences and practices in the context of the energy crisis, IOP Conference Series: Earth and En- vironmental Science 1123 (1) (2022) 012043.doi: 10.1088/1755-1315/1123/1/012043
-
[29]
S. Dey, T. Marzullo, G. Henze, Inverse reinforce- ment learning control for building energy manage- ment, Energy and Buildings 286 (2023) 112941.doi: 10.1016/j.enbuild.2023.112941
-
[30]
M. Liu, M. Guo, Y. Fu, Z. O’Neill, Y. Gao, Expert- guided imitation learning for energy management: Evaluating GAIL’s performance in building control applications, AppliedEnergy372(2024)123753.doi: 10.1016/j.apenergy.2024.123753
-
[31]
H. Zhang, Y. Ding, Z. Tian, An imitation reinforce- ment learning based energy management framework for building air-conditioning systems with chilled wa- ter storage, Energy and Buildings 353 (2026) 116951. doi:10.1016/j.enbuild.2026.116951
-
[32]
G. Besagni, M. Borgarello, The socio-demographic and geographical dimensions of fuel poverty in Italy, Energy Research & Social Science 49 (2019) 192–203. doi:10.1016/j.erss.2018.11.007
-
[33]
L. Campagnolo, E. De Cian, Distributional conse- quences of climate change impacts on residential en- ergy demand across Italian households, Energy Eco- nomics 110 (2022) 106020.doi:10.1016/j.eneco. 2022.106020
doi:10.1016/j.eneco 2022
-
[34]
M. Ferrando, A. Banfi, F. Causone, Changes in en- ergy use profiles derived from electricity smart me- ter readings of residential buildings in Milan before, during and after the COVID-19 main lockdown, Sus- tainable Cities and Society 99 (2023) 104876.doi: 10.1016/j.scs.2023.104876
-
[35]
A. Bahmanyar, A. Estebsari, D. Ernst, The impact of different COVID-19 containment measures on elec- tricity consumption in Europe, Energy Research & Social Science 68 (2020) 101683.doi:10.1016/j. erss.2020.101683
doi:10.1016/j 2020
-
[36]
ISPRA, Tropical Nights (2025)
2025
-
[37]
J. Ballester, M. Quijal-Zamorano, R. F. Méndez Tur- rubiates, F. Pegenaute, F. R. Herrmann, J. M. Robine, X. Basagaña, C. Tonne, J. M. Antó, H. Achebak, Heat-related mortality in Europe during the summer of 2022, Nature Medicine 29 (7) (2023) 1857–1866.doi:10.1038/s41591-023-02419-z
-
[38]
M. Chen, K. T. Sanders, G. A. Ban-Weiss, A new method utilizing smart meter data for identifying the existence of air conditioning in residential homes, En- vironmental Research Letters 14 (9) (2019) 094004. doi:10.1088/1748-9326/ab35a8
-
[39]
Pesaresi, P
M. Pesaresi, P. Politis, GHS-BUILT-S R2023A - GHS built-up surface grid, derived from Sentinel2 composite and Landsat, multitem- poral (1975-2030) (May 2023).doi:10.2905/ 9F06F36F-4B11-47EC-ABB0-4F8B7B1D72EA
1975
-
[40]
Martinez, G
A. Martinez, G. Kakoulaki, P. Florio, P. Poli- tis, DBSM R2025: EU Digital Building Stock Model update including satellite-based attributes, Tech. rep., European Commission, Joint Re- search Centre (JRC) (May 2025).doi:10.2905/ a601a4a8-9289-4fc4-983a-25d54f957f3a
2025
-
[41]
G. Falchetta, A. T. Hammad, Tracking green space along streets of world cities, Environmental Research: Infrastructure and Sustainability 5 (2) (2025) 025011. doi:10.1088/2634-4505/add9c4
-
[42]
H. Jian, Z. Yan, X. Fan, Q. Zhan, C. Xu, W. Bei, J.Xu, M.Huang, X.Du, J.Zhu, Z.Tai, J.Hao, Y.Hu, A high temporal resolution global gridded dataset of human thermal stress metrics, Scientific Data 11 (1) (2024) 1116.doi:10.1038/s41597-024-03966-x
-
[43]
F. M. Bianchi, S. Scardapane, S. Løkse, R. Jenssen, Reservoir Computing Approaches for Representation 16 and Classification of Multivariate Time Series, IEEE Transactions on Neural Networks and Learning Sys- tems 32 (5) (2021) 2169–2179.doi:10.1109/TNNLS. 2020.3001377
doi:10.1109/tnnls 2021
-
[44]
T. G. Kolda, B. W. Bader, Tensor Decompositions and Applications, SIAM Review 51 (3) (2009) 455– 500.doi:10.1137/07070111X
-
[45]
I. T. Jolliffe, Principal Component Analysis, Springer Series in Statistics, Springer-Verlag, New York, 2002. doi:10.1007/b98835
doi:10.1007/b98835 2002
-
[46]
J. H. Ward, Hierarchical Grouping to Optimize an Objective Function, Journal of the American Sta- tistical Association 58 (301) (1963) 236–244.doi: 10.1080/01621459.1963.10500845
-
[47]
P. Virtanen, R. Gommers, T. E. Oliphant, M. Haber- land, T. Reddy, D. Cournapeau, E. Burovski, P. Pe- terson, W. Weckesser, J. Bright, S. J. Van Der Walt, M. Brett, J. Wilson, K. J. Millman, N. Mayorov, A. R. J. Nelson, E. Jones, R. Kern, E. Larson, C. J. Carey, İ. Polat, Y. Feng, E. W. Moore, J. Vander- Plas, D. Laxalde, J. Perktold, R. Cimrman, I. Hen- ...
-
[48]
J. Fu, K. Luo, S. Levine, Learning Robust Re- wards with Adversarial Inverse Reinforcement Learn- ing (Aug. 2018).arXiv:1710.11248,doi:10.48550/ arXiv.1710.11248
Pith/arXiv arXiv 2018
-
[49]
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal Policy Optimization Algorithms (Aug. 2017).arXiv:1707.06347,doi:10.48550/ arXiv.1707.06347
Pith/arXiv arXiv 2017
-
[50]
R. Trauth, M. Kaufeld, M. Geisslinger, J. Betz, Learning and Adapting Behavior of Autonomous Vehicles through Inverse Reinforcement Learning, in: 2023 IEEE Intelligent Vehicles Symposium (IV), 2023, pp. 1–8.doi:10.1109/IV55152.2023. 10186668
-
[51]
A. Likmeta, A. M. Metelli, G. Ramponi, A. Tirin- zoni, M. Giuliani, M. Restelli, Dealing with multi- ple experts and non-stationarity in inverse reinforce- ment learning: An application to real-life problems, Machine Learning 110 (9) (2021) 2541–2576.doi: 10.1007/s10994-020-05939-8
-
[52]
Biewald, Experiment tracking with weights and bi- ases (2020)
L. Biewald, Experiment tracking with weights and bi- ases (2020)
2020
-
[53]
J. Snoek, H. Larochelle, R. P. Adams, Practical Bayesian Optimization of Machine Learning Algo- rithms (2012).doi:10.48550/ARXIV.1206.2944
-
[54]
A. Gleave, M. Taufeeque, J. Rocamonde, E. Jenner, S. H. Wang, S. Toyer, M. Ernestus, N. Belrose, S. Em- mons, S.Russell, Imitation: CleanImitationLearning Implementations (Nov. 2022).arXiv:2211.11972, doi:10.48550/arXiv.2211.11972
-
[55]
M. Towers, A. Kwiatkowski, J. Terry, J. U. Balis, G. D. Cola, T. Deleu, M. Goulão, A. Kallinteris, M. Krimmel, A. KG, R. Perez-Vicente, A. Pierré, S. Schulhoff, J. J. Tai, H. Tan, O. G. Younis, Gymna- sium: A Standard Interface for Reinforcement Learn- ing Environments (Nov. 2025).arXiv:2407.17032, doi:10.48550/arXiv.2407.17032
-
[56]
I. Loshchilov, F. Hutter, Decoupled Weight Decay Regularization (Jan. 2019).arXiv:1711.05101,doi: 10.48550/arXiv.1711.05101
-
[57]
T. Miyato, T. Kataoka, M. Koyama, Y. Yoshida, Spectral Normalization for Generative Adversarial Networks (Feb. 2018).arXiv:1802.05957,doi:10. 48550/arXiv.1802.05957
Pith/arXiv arXiv 2018
-
[58]
S. Ioffe, C. Szegedy, Batch Normalization: Accel- erating Deep Network Training by Reducing Inter- nal Covariate Shift (Mar. 2015).arXiv:1502.03167, doi:10.48550/arXiv.1502.03167
-
[59]
L. Engstrom, A. Ilyas, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, A. Madry, Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO (May 2020).arXiv:2005.12729, doi:10.48550/arXiv.2005.12729
-
[60]
P. Isola, J.-Y. Zhu, T. Zhou, A. A. Efros, Image- to-Image Translation with Conditional Adversarial Networks (Nov. 2018).arXiv:1611.07004,doi:10. 48550/arXiv.1611.07004. 17
Pith/arXiv arXiv 2018
-
[62]
Salimans, I
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, X. Chen, X. Chen, Improved Techniques for Training GANs, in: Advances in Neural Informa- tion Processing Systems, Vol. 29, Curran Associates, Inc., 2016
2016
-
[63]
Pascanu, T
R. Pascanu, T. Mikolov, Y. Bengio, On the difficulty of training recurrent neural networks, in: Proceed- ings of the 30th International Conference on Machine Learning, PMLR, 2013, pp. 1310–1318
2013
-
[64]
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, K. He, Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour (Apr. 2018).arXiv:1706.02677,doi: 10.48550/arXiv.1706.02677
-
[65]
M. Arjovsky, L. Bottou, Towards Principled Methods for Training Generative Adversarial Networks (Jan. 2017).arXiv:1701.04862,doi:10.48550/arXiv. 1701.04862
-
[66]
Istat, Viaggi e vacanze in Italia e all’estero, Statistical report, Istituto Nazionale di Statistica (Istat), Roma, Italia (2025)
2025
-
[67]
B. Stikvoort, A. Nilsson, C. Bartusch, V. van Zoest, In the rhythm of the home: How does increased home occupancy affect residential electricity consumption?, Energy Research & Social Science 123 (2025) 104032. doi:10.1016/j.erss.2025.104032
-
[68]
Y. Fan, J. Wang, N. Obradovich, S. Zheng, Intra- day adaptation to extreme temperatures in outdoor activity, Scientific Reports 13 (1) (2023) 473.doi: 10.1038/s41598-022-26928-y
-
[69]
I. Batur, V. O. Alhassan, M. V. Chester, S. E. Polzin, C. Chen, C. R. Bhat, R. M. Pendyala, Understanding how extreme heat impacts human activity-mobility and time use patterns, Transportation Research Part D: Transport and Environment 136 (2024) 104431. doi:10.1016/j.trd.2024.104431. 18
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.