{"id":"0d81762b-3ae2-4766-b46f-10ebc4da9161","arxiv_id":"1908.08176","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A regression-plus-clustering pipeline benchmarks residential AC rooms on predicted power at equalized weather and settings, tested on 44 Singapore rooms.","lead":"The paper proposes a data-driven method to benchmark the air-conditioning energy performance of individual residential rooms, using per-room regression models and clustering to compare rooms under identical weather and settings. It is a practical approach for giving households fair energy-efficiency feedback without simulation or manual calibration.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Uniform noisy-factor values are selected only by marginal overlap (Section 4.3.2); for rooms with sparse or shifted data, predictions at those values are extrapolations, so the 'valid and fair' claim lacks direct support.","rationale":"I agree with the reader's identified weakest assumption: it is the point on which the central claim depends. The proposed pipeline could still be a reasonable applied contribution, but the paper's own Section 4.3.1 admits extrapolation risk, and Step 3.2's marginal-overlap criterion is not sufficient to control it. The concern is internally grounded in the described procedure and is testable. I do not see a reason to move beyond conditional acceptance: the approach is clearly described and the case study is informative, but the fair-benchmarking claim should be conditional on demonstrating that uniform-point predictions lie within each room's validated prediction region. The reader's verdict of CONDITIONAL is therefore appropriate, and my independent stress-test identifies the same load-bearing weakness rather than a different one.","tokens_in":23368,"tokens_out":3393,"duration_ms":36882,"concrete_test":"For each room and its cluster's uniform noisy-factor vector, define a neighborhood in the normalized feature space (e.g., Mahalanobis distance <= 1 using that room's training covariance, or a kernel bandwidth around the point). Count historical training segments inside that neighborhood and compute actual-versus-predicted MAPE on those segments using a model trained without them. Flag rooms with fewer than a minimum count (e.g., 5) as extrapolating. Then recompute the benchmarking scores after excluding flagged rooms or replacing their predictions with a more conservative estimate; if the mean absolute change in scores exceeds 0.1 or any best-room ranking changes, the fair-benchmarking claim is not supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the benchmarking scores are 'valid and fair' and reveal 'the real AC energy performance' (Section 5.4.2, Conclusion). This requires each room's regression model to be accurate at the uniform noisy-factor vector used for comparison inside its cluster. The uniform values are selected in Section 4.3.2 by comparing marginal percentile ranges of historical values and taking the center of the overlap; the procedure never verifies that the joint seven-dimensional feature vector lies inside each room's training support. Section 4.3.1 itself acknowledges that 'predictive model usually performs worse on input that it has not seen during the training.' With nseg_min = 20 (Section 5.1) and rooms split by tenant, a room can have sparse or narrowly distributed data (e.g., R14-3's set-point distribution in Fig. 5d). For such a room, the uniform vector can sit at the edge of, or outside, its observed region even when every marginal range overlaps. Its predicted power is then an extrapolation, and the KDE residual model from Section 4.2.2, fitted on cross-validation residuals from the training region, is not representative at that point. Since scores are ratios to the lowest predicted power in the cluster, one extrapolated room can become the benchmark or be unfairly ranked worst. The case-study validation (Types A and B, Section 5.4.2) compares only historical-score variants and does not test accuracy at the uniform operating point, so it cannot detect this bias.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a data-driven, peer-to-peer benchmarking method for air-conditioning energy performance of residential rooms. For each room, a regression model is trained to predict average AC power from seven segment-wise factors (ambient temperature, humidity, solar irradiance, initial room temperature, average room temperature, segment duration, and temperature set point). Rooms are then clustered by room area and median set point using k-means; within each cluster, a uniform vector of the seven segment-wise factors is derived from overlapping historical percentile ranges, and each room's fitted model is evaluated at that vector to produce a stochastic predicted power. The benchmarking score is the ratio of the cluster's lowest predicted power to the room's predicted power. A case study on 44 rooms reports an average cross-validated MAPE of 14.9%, seven clusters selected by Silhouette value, and a comparison between the proposed scores and two historical-data baselines (Type A and Type B). The authors conclude that the approach eliminates influences of room area, weather, and AC settings and therefore is 'valid and fair.' The core technical pipeline is coherent and the case study is informative, but the validation does not establish the central fairness claim: it shows only that scores change after equalization, not that the new scores are correct.","tokens_in":23670,"tokens_out":5812,"duration_ms":56761,"significance":"Residential room-level AC benchmarking is a genuine and understudied problem, and the proposed combination of per-room regression, clustering, and noisy-factor equalization is a sensible and scalable design. If the equalization were validated against an independent ground truth, the method would be practically useful for energy feedback programs and could be transferred to similar settings. The paper's literature review organized around benchmarking steps is a useful contribution, and the empirical study provides a transparent account of model selection, including cross-validation accuracy and computational cost. However, the central claims of validity and fairness are not adequately supported: the validation compares only internally derived score variants, and the extrapolation and orientation concerns mean the paper currently demonstrates consistency of a procedure rather than correctness of a benchmark.","major_comments":[{"comment":"The central claim that the proposed benchmarking scores are 'valid and fair' and reveal 'the real AC energy performance' is not supported by the validation. Type A and Type B scores are both computed from the same historical power data used to build the room models; the finding that the proposed scores differ from Type B after equalization (mean absolute difference 0.1378) demonstrates only that the procedure changes the ranking, not that the new ranking is correct. If the regression models are biased at the evaluation point, the equalized scores inherit that bias. An independent check is needed, for example a field experiment with controlled operation, comparison against manufacturer EER or physical measurements, or a synthetic-data study in which the true energy performance is known.","section":"Section 5.4.2 and Conclusion"},{"comment":"The uniform seven-dimensional noisy-factor vector is selected using only marginal percentile overlaps of each factor's historical range. This does not guarantee that the vector lies inside the joint training support of every room's regression model. The paper itself states in Section 4.3.1 that a predictive model 'usually performs worse on input that it has not seen during the training,' yet Step 3.2 checks only whether historical percentile ranges overlap marginally. For a room with sparse or tightly distributed data, e.g., R14-3's set-point distribution in Fig. 5d, the chosen uniform value can be an extrapolation even when every marginal range overlaps. The predicted power is then unreliable, and because the score is a ratio to the minimum predicted power in the cluster, one extrapolated room can set the benchmark. The authors should add a coverage check (e.g., the density of each room's training data at the uniform vector, or the Mahalanobis distance from the training mean) and either restrict to rooms or clusters where coverage is adequate or quantify the uncertainty in the score.","section":"Section 4.3.2 (Step 3.2) with Section 4.3.1"},{"comment":"The noisy-factor list omits room orientation and window/solar aperture properties, although Table 2 reports that the studied rooms have four different orientations. In the thermal model of Eq. (8), solar gain is represented only by global solar irradiance (psi) times a conversion ratio; the effect of orientation on the amount of solar radiation reaching the room is not captured. Since orientation is an objective building characteristic outside the user's control, two rooms in the same cluster with east and west orientations can receive different solar loads at the same measured irradiance, and the benchmarking comparison then penalizes the west-facing room rather than its energy performance. The fairness claim requires either adding orientation to the clustering features and a per-orientation solar model, or an explicit assumption about uniformity of orientation within clusters.","section":"Section 3.3 and Table 2"}],"minor_comments":[{"comment":"The caption contains a typo: 'sevens rooms' should be 'seven rooms,' and 'overlaping' should be 'overlapping.'","section":"Section 5.4.1, Fig. 5 caption"},{"comment":"The sentence beginning 'Tn the filtering process' should read 'In the filtering process.'","section":"Section 5.1"},{"comment":"The phrase 'to balane between the representativeness and computational cost' contains a typo; 'balane' should be 'balance.'","section":"Section 5.4.1"},{"comment":"The sentence 'The computation of the uniform values of the seven noisy factors for all rooms in the same cluster is done factor by factor, but in the same manner summarized in the following' is grammatically incomplete; rephrase to make the enumeration explicit.","section":"Section 4.3.2"},{"comment":"The arrows in Fig. 3 indicate how hyper-parameter settings affect model performance, but this is not explained in either the figure caption or the main text; add a clarifying sentence.","section":"Figure 3 and Table 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is potentially publishable after substantial revision. The most important issue is the mismatch between the paper's strong conclusion ('valid and fair,' 'real AC energy performance') and the internal-only validation. If the authors can add at least one independent validation or carefully limit the claims to 'equalization of selected noisy factors' rather than 'valid and fair,' the paper would be much stronger. The extrapolation and orientation concerns also need to be addressed explicitly, either with additional diagnostics or by narrowing the scope of the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plain English: this paper fills a genuine gap, but the central claim is not fully backed. The pipeline is sensible: per-room regression models, clustering on area and set point, then equalizing the other seven noisy factors by evaluating all models at a uniform input vector. That combination is new in the residential AC benchmarking literature, and the case study on 44 real rooms, including tenant-split virtual rooms, is a solid empirical exercise. The model comparison across 13 structures is careful, and the cross-validation accuracy is honestly reported.\n\nWhere it gets soft: the 'valid and fair' conclusion is asserted from an internal comparison. The validation shows that scores change after equalization relative to Type A and Type B scores, but all three are computed from the same historical data. There is no independent measure of AC efficiency—no controlled test, no known-efficiency benchmark room, no simulation ground truth—so the equalized scores could be consistently wrong and the demonstration would still look convincing.\n\nThe stress-test concern about extrapolation is real. The uniform noisy-factor vector is chosen as the center of overlapping marginal percentile ranges per factor. That guarantees each factor overlaps marginally, not that the joint seven-dimensional vector lies inside each room's training support. Section 4.3.1 admits models perform worse on unseen inputs, but Step 3.2 never checks the uniform vector against each room's actual data density. With nseg_min = 20, some rooms have thin datasets, and a room with sparse data near the edge of the overlap can be evaluated largely by extrapolation. Since scores are ratios to the lowest predicted power, one extrapolated room can distort the whole cluster ranking.\n\nA third issue: building orientation appears in Table 2 but is never used as a noisy factor. Solar irradiance is measured, but a west-facing room will absorb more heat than a north-facing room at the same irradiance level. If orientations are not balanced within clusters, the scores will reflect envelope differences, not just AC performance. This is not fatal, but it is an unexamined assumption worth flagging.\n\nWho gets value: researchers working on residential energy benchmarking, smart-meter feedback, or data-driven performance assessment. The method is a useful template and the limitations are honestly discussed (e.g., singleton clusters). It deserves a serious referee, but the language about being 'valid and fair' and revealing 'the real AC energy performance' should be tempered in revision, and the extrapolation issue should be addressed—perhaps by checking the uniform vector against each room's joint data support or by adding a synthetic validation experiment.","headline":"Competent applied paper with a real niche, but the 'valid and fair' conclusion is stronger than the internal validation supports.","tokens_in":24149,"tokens_out":1909,"would_cite":false,"duration_ms":23043,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that air-conditioning energy performance of residential rooms can be benchmarked fairly by building a regression model for each room, clustering comparable rooms, and comparing predicted power under identical weather and…","keywords":["energy performance benchmarking","air conditioning","residential rooms","regression","support vector regression","clustering","noisy factor equalization","peer-performance benchmark"],"falsifier":"Take one cluster and the uniform noisy-factor values computed for it. For each room, count the historical operation segments whose features fall within a small neighborhood of that point, for example within one standard deviation of each chosen value. If the room ranked worst has few or no segments there while the room ranked best has many, its predicted power is an extrapolation. A direct test: retrain each room's model only on segments in that neighborhood and recompute the scores; if the rankings change materially, the fairness claim fails.","tokens_in":23174,"feed_emoji":"❄️","tokens_out":5463,"duration_ms":53538,"temperature":0.7,"pith_summary":"Air conditioning is a large and growing share of global electricity use, and the paper argues that the residential part of it can be benchmarked fairly, room by room. The central claim is that a data-driven pipeline, per-room regression models, clustering of comparable rooms, and comparison under one identical set of weather and control conditions, eliminates the influence of room area, weather, and AC settings, so the resulting scores reveal the actual AC energy performance hidden beneath historical data. If this is right, residents can be given a peer-relative score that points to poor maintenance, envelope leakage, or wasteful habits without blaming them for wanting a cooler room or living in a larger one. The paper demonstrates the approach on 44 real rooms, reporting average prediction accuracy of 85.1% in cross-validation and seven comparable clusters.","feed_headline":"Fair AC benchmark erases room size, weather, and settings","feed_subtitle":"Each room is scored against peers under identical conditions, exposing the real performance hidden in historical data.","key_machinery":"The load-bearing mechanism is the pair of a per-room regression model and a cluster. For each room, a regression model $F(\\bar{T}_a, \\bar{H}_a, \\bar{p}_{si}, T_{ri}, \\bar{T}_r, t_{seg}, \\bar{T}_{set})$ predicts average AC power $\\bar{p}_{ac}$ from segment-wise weather, temperature, duration, and set-point inputs; its percentage residual is modeled by kernel density estimation and sampled so each room's representative energy performance index is stochastic. Rooms are grouped by k-means clustering on room area and median set point, and within each cluster a uniform value set is computed from overlapping percentile ranges of historical data. The benchmark is the lowest stochastic predicted power in the cluster, and the score is $\\eta = \\hat{p}_{ac}^* / \\hat{p}_{ac}$. The regression handles factors that vary segment to segment, the clustering handles the two factors that define comparability, and the uniform value set equalizes the rest.","core_discovery":"The paper's discovery claim is that benchmarking the AC energy performance of residential rooms is feasible as a fully data-driven peer comparison: instead of averaging historical power use, each room gets a regression model of average power per AC operation segment as a function of seven noisy factors (ambient temperature, humidity, solar irradiance, initial room temperature, average room temperature, segment duration, and temperature set point), and rooms are clustered by room area and median set point. Within each cluster, one uniform set of noisy-factor values is chosen from the overlap of the rooms' historical ranges; every room's predicted power at that point, made stochastic by sampling the model's percentage residual, is compared with the lowest predicted power in the cluster. The ratio of best to room becomes the benchmarking score. Using 44 rooms, the paper shows that this procedure changes rankings relative to raw historical comparisons, including reversing the order of two tenants of the same physical room, and concludes that the scores are valid and fair because all eight selected noisy factors have been equalized.","pith_inferences":["The fairness claim is conditional on model accuracy at the uniform query point; a natural robustness check is to restrict scoring to rooms whose training data actually cover that point, and this could reshuffle rankings in clusters with sparse overlap.","Because the score is a ratio to the best room in the same cluster, it measures relative performance, not absolute efficiency; a cluster of uniformly wasteful rooms would still produce scores near one, so the method identifies outliers rather than an efficiency target.","The same regression-plus-equalization template transfers to other decentralized energy systems such as heat pumps, water heaters, or lighting, where energy use depends on user settings and ambient conditions, with the noisy-factor list replaced by domain equivalents.","The tenant-reversal result in the case study suggests the method separates user behavior from the physical room, but confirming that separation as a measure of hardware or maintenance quality would need independent envelope and equipment data, which the paper leaves to future work."],"forward_implications":["Within a cluster, benchmarking scores reflect only differences the paper attributes to AC system quality, maintenance, and user behavior, because all eight noisy factors are equalized.","Rooms that historically operated under different weather and settings can be compared directly, as long as their historical ranges overlap.","The approach runs without manual benchmark selection or simulation calibration, using only power, temperature, weather, and room-area data, so it scales to larger sets of rooms.","The trained models can be reused to simulate how room area and temperature set point affect AC power, turning the benchmark into an analytical tool.","A room that ends up alone in its cluster receives no meaningful peer score; the paper proposes supplementing with previous-performance benchmarking for such cases."],"supporting_citations":[{"why":"Supplies the core premise that objective noisy factors must be eliminated from energy comparisons so benchmarking results are fair.","marker":"[19]"},{"why":"Provides the three-way taxonomy of previous-performance, intended-performance, and peer-performance benchmarks that justifies the peer-performance choice.","marker":"[21]"},{"why":"Shows that complex cooling-load and efficiency indices require detailed system data, which motivates the scalable direct-power EPI used here.","marker":"[16]"},{"why":"Supplies the clustering-based analysis benchmark procedure that the proposed method extends by adding per-room regression equalization.","marker":"[35]"},{"why":"Supports the use of intelligent clustering to group comparable entities before within-group benchmarking.","marker":"[39]"},{"why":"Gives the inverse black-box regression modeling approach used to predict cooling power from operating conditions.","marker":"[40]"},{"why":"Provides the Silhouette criterion used to select the number of clusters in k-means.","marker":"[41]"},{"why":"Supplies the instrumented AC data used in the case study, including power, operation settings, and indoor temperature measurements.","marker":"[42]"}],"fun_headline_variants":["Regression and clustering strip bias from AC efficiency scores","Fair AC scores: regression models erase room size and weather","Data-driven AC benchmarking neutralizes area, climate, settings","Clustering equalizes conditions to rank AC performance fairly","AC energy ratings get fair overhaul with regression clustering"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The approach assumes that the single uniform set of weather and setting values used to compare rooms sits inside the reliable prediction region of every room's regression model; the overlap check covers historical percentile ranges, not the accuracy of each model at the chosen point.","fun_headline_variants_meta":{"raw":{"variants":["Regression and clustering strip bias from AC efficiency scores","Fair AC scores: regression models erase room size and weather","Data-driven AC benchmarking neutralizes area, climate, settings","Clustering equalizes conditions to rank AC performance fairly","AC energy ratings get fair overhaul with regression clustering"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000577,"raw_usage":{"total_tokens":2765,"prompt_tokens":1032,"completion_tokens":1733,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":1657}},"tokens_in":648,"tokens_out":1733,"duration_ms":13206,"temperature":1.0,"reasoning_tokens":1657,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:47:18.949704+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one cluster and the uniform noisy-factor values computed for it. For each room, count the historical operation segments whose features fall within a small neighborhood of that point, for example within one standard deviation of each chosen value. If the room ranked worst has few or no segments there while the room ranked best has many, its predicted power is an extrapolation. A direct test: retrain each room's model only on segments in that neighborhood and recompute the scores; if the rankings change materially, the fairness claim fails.","supporting_citations":[{"cited_title":"Chung, Review of building energy-use performance benchmarking methodologies, Applied Energy 88 (5) (2011) 1470–1479","cited_arxiv_id":null,"evidence_quote":"Supplies the core premise that objective noisy factors must be eliminated from energy comparisons so benchmarking results are fair."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the three-way taxonomy of previous-performance, intended-performance, and peer-performance benchmarks that justifies the peer-performance choice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that complex cooling-load and efficiency indices require detailed system data, which motivates the scalable direct-power EPI used here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the clustering-based analysis benchmark procedure that the proposed method extends by adding per-room regression equalization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the use of intelligent clustering to group comparable entities before within-group benchmarking."},{"cited_title":"Gunay, W","cited_arxiv_id":null,"evidence_quote":"Gives the inverse black-box regression modeling approach used to predict cooling power from operating conditions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Silhouette criterion used to select the number of clusters in k-means."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the instrumented AC data used in the case study, including power, operation settings, and indoor temperature measurements."}],"review_version":1}