{"id":"b469c0e9-aaf7-447a-b5df-41e7a83e4920","arxiv_id":"2504.13483","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"NPIL, a latent tensor factorization model that rewrites the SGD error with a nonlinear PID controller and PSO-tuned gains, recovers missing smart-meter readings faster and more accurately than five comparison models on iAWE and UK-DALE.","lead":"This paper combines a nonlinear PID controller with tensor factorization, tuned by particle swarm optimization, to fill in missing power-measurement data from home energy monitors. It reports faster convergence and slightly better accuracy than five older tensor models on two real home-power datasets.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'state-of-the-art' claim is not yet established because the comparison set contains no NILM-specific imputer and only random missing masks are tested.","rationale":"The reader's CONDITIONAL verdict is well calibrated. The numerical tables are internally consistent, the wall-clock advantage over the tested tensor baselines is real, and the paper is candid about building on published PID-LFT components. My stress-test identifies the same core limitation from a slightly different angle: the load-bearing premise is not the stability of the NPID update but the external validity of the comparison. Since no NILM-specific imputer and no structured-missing benchmark appear, the abstract's 'surpasses state-of-the-art' claim cannot be verified from this paper alone. That does not mean the model is wrong; it means the central claim needs either an in-domain baseline comparison or a narrowed statement. A single structured-missing/in-domain-baseline experiment would settle whether the concern lands, so the conditional verdict is retained unchanged.","tokens_in":17085,"tokens_out":5705,"duration_ms":52903,"concrete_test":"Reproduce the iAWE and UK-DALE experiments with (i) the same random masks and (ii) structured masks that delete contiguous blocks of channels/time (simulating sensor dropout), and add at least three NILM-relevant imputation baselines: a standard interpolator/kNN, a matrix/tensor completion method (e.g., softImpute or CP-WOPT), and a sequential deep imputer (e.g., BRITS or a Transformer/LSTM). If NPIL does not outperform the best such baseline in RMSE/MAE and wall-clock under both mask types, the 'state-of-the-art' claim should be narrowed to the tested tensor-factorization family.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that NPIL 'surpasses state-of-the-art models' for missing NILM data. For that claim to hold, the comparison class must include methods that actually represent the NILM imputation state of the art, and the missingness regime should resemble NILM failures. Section IV.A creates only 'random missing scenarios' (95%/90%/85% missing), and Section IV.B compares NPIL with M1-M5, which are tensor-factorization/QoS/dynamic-network/fake-news models ([66], [74]-[77]); none is a NILM-specific imputer, and no deep-learning imputation method (LSTM, Transformer, or BRITS) is included. Thus the reported gains (e.g., 5.19% RMSE and 7.04% MAE over M3 on iAWE at 15%) establish only that NPIL beats a set of adjacent-domain baselines under random masks. The motivating failure mode, sensor failure, typically produces structured missing blocks; the paper provides no evidence that the PSO-tuned NPID gains remain stable or accurate there. The concern is scope, not internal inconsistency: the headline claim is stronger than the experiment design supports.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes NPIL, a nonlinear PID-incorporated latent factorization of tensors model for recovering missing NILM data. The model replaces the raw instantaneous error in an SGD-based tensor factorization with a nonlinear PID-adjusted error that includes an integral term over past errors and sech/exp gain nonlinearities, and it tunes the nine PID gains with particle swarm optimization. Experiments on iAWE and UK-DALE at 95%, 90%, and 85% missing rates compare NPIL with five baselines and report lower RMSE and MAE, faster convergence, and lower time cost. The claims are empirical: the central contribution is a modified SGD update with PSO-based gain adaptation, and the evaluation is a benchmark study against adjacent-domain tensor and QoS models under random missingness.","tokens_in":17186,"tokens_out":5739,"duration_ms":50713,"significance":"If the empirical claims hold, NPIL is a practically useful method for random-missingness high-sparsity tensor imputation: it reports consistent RMSE/MAE improvements over the included baselines, degrades only mildly as density drops from 15% to 5%, and converges in very few iterations. The paper's internal arithmetic is consistent, and the use of two public datasets with 15 repeated runs is a strength. However, the headline claim of surpassing 'state-of-the-art models' for NILM missing-data recovery is not yet established: the baseline set is narrow and the missingness regime is limited. The novelty relative to prior nonlinear PID-incorporated SGD for latent factor analysis also needs clearer positioning. With rescoped claims or additional experiments, the core idea would be a reasonable contribution to tensor-based data imputation.","major_comments":[{"comment":"The headline claim that NPIL 'surpasses state-of-the-art models' for NILM missing-data recovery is not supported by the comparison design. Section IV.A creates only random missing scenarios, whereas the motivating failure mode (sensor failure, Section I) typically produces structured or block-wise missingness. Section IV.B compares NPIL with five baselines (M1-M5) from tensor factorization, QoS prediction, and fake-news detection; none is a NILM-specific imputer, and no deep sequence imputation method (e.g., LSTM-, Transformer-, or BRITS-based) is included. The reported gains therefore establish superiority only over a restricted baseline set under random masks. Please either add NILM-specific and structured-missingness experiments or explicitly scope the claim to the tested regime.","section":"Section IV.A and IV.B (TABLE III, baseline list)"},{"comment":"The convergence and accuracy claims rest on the nonlinear PID-adjusted error update, but the paper provides no convergence or stability analysis for the modified SGD recursion, which includes an accumulated integral term and sech/exp gain nonlinearities. The '5 iterations' convergence observation (Section IV.B b) is empirical and demonstrated only for the PSO-searched gain ranges in TABLE II. Because the integral accumulation can in principle cause oscillation or divergence under other missingness patterns, densities, or datasets, please provide either a formal stability/convergence argument or an empirical robustness study over structured missingness and gain perturbations.","section":"Section III.A, Eqs. (10)-(11)"},{"comment":"The convergence-rate comparison is ambiguous because the iteration counts are not directly comparable across models. The text states that each NPIL iteration includes Q PSO sub-iterations, so a comparison in terms of iterations to convergence (5 iterations for M6 on D2-15%) conflates algorithmic epochs with wall-clock time. The time-cost results in TABLE IV are the appropriate basis for the efficiency claim and should be presented as the primary evidence; otherwise, please report wall-clock convergence curves for all models.","section":"Section IV.B b) and TABLE IV"}],"minor_comments":[{"comment":"The sentence '26.49%lower than M5’s 13.89s, M6 incurs significant time costs' is garbled and self-contradictory; please rephrase and correct the table reference (TABLE II vs TABLE IV).","section":"Section IV.B c)"},{"comment":"The text refers to 'Fig. 4' and 'TABLE Ⅱ' for training curves and time costs, but the figure is captioned Fig. 3 and time costs are in TABLE IV.","section":"Section IV.A and Fig. 3"},{"comment":"The fitness function in Eq. (13) is typeset unclearly, and the quantities p=u=0.5 are not defined; please rewrite (13) in standard notation and define all symbols.","section":"Eq. (13)"},{"comment":"The nonlinear gain forms are introduced without justification; they should be described as heuristic gain-shaping functions, and any boundedness or positivity constraints needed for stable updates should be stated.","section":"Section II.B, Eq. (4)"},{"comment":"No statistical significance tests are reported for the differences in TABLE III; given the small margins over M3, please add a paired test or confidence intervals.","section":"TABLE III"},{"comment":"The explanation of the 8:1:1 split into training, test, and validation sets is ambiguous: it should state whether the split is over observed entries or over the full tensor, and how the random missing masks interact with the split.","section":"Section IV.A"}],"recommendation":"major_revision","confidential_remarks":"The paper is in scope for a tensor-imputation journal, but the novelty relative to [54]—which already proposes nonlinear PID-incorporated SGD for latent factor analysis—appears incremental; the new elements are PSO-based gain adaptation and the NILM application. The authors should be required to state the technical difference more explicitly. Also, the absence of NILM-specific baselines may be acceptable if the claims are rescoped, but the current framing is likely to draw criticism from readers in the NILM community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid applied combination paper, not a breakthrough. The method—nonlinear PID-adjusted errors inside CP tensor factorization, with PSO tuning the nine gains—is assembled from published pieces (linear PID in [47],[48], nonlinear PID for matrices in [54], PSO in [34],[73],[80]), and the paper is upfront about that lineage. What it adds for NILM is the tensor-level nonlinear PID formulation with self-adaptive gains, and the experiments are internally consistent: NPIL gets the lowest RMSE/MAE across all densities on both iAWE and UK-DALE, and the numbers hang together (the 2.67% vs. 60.59% degradation comparison is arithmetic, not hand-waving). The wall-clock gains over M2/M3 are large, and the convergence curves look like genuine results. If you work on missing-data recovery for smart-meter tensors, this is a plausible candidate to try.\n\nThe soft spots are mostly about scope and reproducibility. The comparison set M1–M5 is all adjacent-domain tensor/QoS/fake-news models; no NILM-specific imputer and no deep-learning baseline (LSTM, BRITS, transformer) is included, so the abstract's 'state-of-the-art' claim is not established. Missing data is random only, while the motivating failure—sensor failure—typically produces structured block missingness. The convergence-rate claim counts iterations, not wall-clock per full epoch, and each iteration bundles Q=5 PSO sub-iterations; the paper admits this, but it still inflates the apparent speed advantage. The nine NPID gains are tuned on validation splits of the test datasets, with η and λ grid-searched on one of the 15 runs; that is per-dataset tuning, so reported numbers would likely shift on new data. No code is shipped, and some implementation details (velocity bounds, gain formulas) are garbled in the text. There is no convergence or stability analysis for the integral accumulation; the '5 iterations' result is empirical for the searched gain ranges.\n\nNone of this kills the paper. The central claim, as bounded by the actual experiment suite, is defensible: NPIL beats those five baselines under random masks. It just does not yet support the broader 'state-of-the-art' sentence. I would send this to a serious referee—it is a competent applied contribution with a clear mechanism and credible internal evidence—but the referee should push for in-domain baselines, structured missingness experiments, and a code/data release. The paper deserves review, not desk rejection, but it needs substantive revision to justify the headline.","headline":"A competent combination of nonlinear PID and PSO in CP tensor factorization for NILM imputation; the internal evidence is credible, but the 'state-of-the-art' claim outruns an experiment suite with only random masks and adjacent-domain baselines.","tokens_in":17902,"tokens_out":2055,"would_cite":false,"duration_ms":18196,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Nonlinear PID error feedback makes tensor factorization fill missing smart-meter readings faster and more accurately than five baselines.","keywords":["non-intrusive load monitoring","missing data recovery","latent factorization of tensors","nonlinear PID controller","particle swarm optimization","tensor completion","smart grid","stochastic gradient descent"],"falsifier":"Run NPIL on UK-DALE at 15% density with entries missing in contiguous blocks, for example all three measurement channels absent for several consecutive hours to mimic sensor dropout, and compare against M3; if RMSE degrades to baseline levels or the training curve oscillates, the claim that the method recovers missing NILM data under sensor failure over-generalizes.","tokens_in":16660,"feed_emoji":"⚡","tokens_out":9986,"duration_ms":84357,"temperature":0.7,"pith_summary":"This paper tries to establish that missing entries in non-intrusive load monitoring data can be recovered more accurately and far faster if the standard tensor-factorization update is driven by a nonlinear PID controller rather than by the raw current error. The proposed NPIL model arranges household power measurements as a third-order tensor, factors it by canonical polyadic decomposition, and in each stochastic gradient descent step replaces the instantaneous residual with a PID-style error that combines proportional, accumulated-integral, and derivative terms under nonlinear gains. Particle swarm optimization tunes the nine controller gains on a validation split before training. On the iAWE and UK-DALE datasets at 5%, 10%, and 15% observed density, the paper reports lower RMSE and MAE than five comparison baselines, with RMSE converging in about five iterations on UK-DALE at 15% density and total runtime near ten seconds on iAWE. This matters because fast, accurate imputation of missing smart-meter readings is a practical precondition for reliable demand response and appliance-level monitoring.","feed_headline":"Five iterations recover missing smart-meter readings","feed_subtitle":"A tensor model with PID-style error feedback beats five baselines on iAWE and UK-DALE power data.","key_machinery":"The central machinery is a nonlinear PID controller used as an error-refinement rule inside stochastic gradient descent for CP tensor factorization. In the objective of Eq. (8), the raw learning residual $e_{ijk}=y_{ijk}-\\hat{y}_{ijk}$ is replaced by $\\tilde{e}_{ijk}$ from Eq. (10), so the update rules in Eq. (11) for the latent factor matrices and bias vectors carry past information through the integral term, anticipate changes through the derivative term, and modulate both through sech and exponential gain functions. The nonlinear gains let the controller respond differently to large and small errors, and particle swarm optimization pre-tunes the nine gain parameters on a validation split rather than by hand. This error-refinement step is what the paper identifies as the cause of the faster convergence and improved accuracy reported in Section IV.","core_discovery":"The paper's central claim is that replacing the instantaneous SGD error with a nonlinear PID-adjusted error yields a latent tensor factorization model, called NPIL, that converges faster and predicts missing non-intrusive load monitoring entries more accurately than five comparison baselines. The adjusted error $\\tilde{e}_{ijk}$ in Eq. (10) combines a proportional term $K_{P1}+K_{P2}(1-\\operatorname{sech}(K_{P3}e))$, an integral term $K_{I1}\\operatorname{sech}(K_{I2}e)\\sum_{h=0}^{t-1}e^{(h)}$, and a derivative term scaled by $K_{D1}+K_{D2}/(1+\\exp(K_{D3}e+K_{D4}))$, and these gains are pre-tuned by particle swarm optimization on a validation split. On iAWE at 15% density, NPIL reaches RMSE 0.0256 and MAE 0.0185, which the paper reports as 5.19% and 7.04% lower than the best baseline M3 and 31.73% and 28.29% lower than M4. On UK-DALE at 15% density, RMSE converges in roughly five iterations and MAE in roughly three, and time cost on iAWE is about 10.21 seconds. The same qualitative conclusion is drawn at 10% and 5% density, where NPIL's error grows only mildly while comparison baselines degrade sharply.","pith_inferences":["An untested extension is to apply the same nonlinear PID error substitution to other tensor decomposition families, such as Tucker or graph-regularized factorizations, since the error-refinement rule does not depend on the canonical polyadic decomposition form; the paper only demonstrates it inside CP-based tensor factorization.","The experiments create random missing entries, whereas the stated motivation of sensor failure typically produces contiguous blocks of missing readings; whether the integral accumulation helps or destabilizes under block-missing structure is an open empirical question the paper does not address.","If the convergence speed persists in an online setting, the model could be embedded in a streaming smart-meter pipeline that imputes gaps as they appear; the paper reports batch experiments only, so this is an extrapolation."],"forward_implications":["At 5% observed density, NPIL's RMSE on iAWE rises only from 0.0256 to 0.0263, while M5's RMSE rises from 0.0342 to 0.0868, so the method remains accurate under extreme sparsity.","Because RMSE converges in about five iterations on UK-DALE at 15% density, retraining on newly arriving smart-meter data can be nearly real-time rather than requiring hundreds of iterations.","On iAWE at 15% density, NPIL's reported runtime of 10.21 seconds is about 97.7% less than M2, 95.5% less than M3, 73.2% less than M4, and 26.5% less than M5.","The same PSO-tuned gain ranges work across two datasets and three missingness levels without manual retuning, so the gain adaptation is a practical part of the method, not a one-off calibration."],"supporting_citations":[{"why":"Supplies the nonlinear PID-incorporated adaptive SGD rule that NPIL applies to tensor factorization.","marker":"[54]"},{"why":"Provides the PSO-incorporated latent factor analysis scheme used to adapt NPIL's gain parameters.","marker":"[80]"},{"why":"Defines M1, the biased non-negative LFT baseline whose RMSE and MAE NPIL must beat.","marker":"[66]"},{"why":"Defines M2, the tensor-based QoS prediction baseline used in the runtime comparison.","marker":"[74]"},{"why":"Defines M3, the CP decomposition baseline reported as the strongest non-NPIL accuracy competitor.","marker":"[75]"},{"why":"Defines M4, the sparsity-and-graph-regularized tensor decomposition baseline.","marker":"[76]"},{"why":"Defines M5, the Cauchy-loss outlier-resilient QoS prediction baseline.","marker":"[77]"},{"why":"Supplies the iAWE dataset used for all D1 missing-data experiments.","marker":"[81]"},{"why":"Supplies the UK-DALE dataset used for all D2 missing-data experiments.","marker":"[82]"}],"fun_headline_variants":["PID-tuned tensor model fills smart-meter gaps faster","Nonlinear PID speeds up missing-data tensor recovery","Five iterations to accurate NILM data recovery","PSO-tuned PID boosts tensor factorization convergence","Missing smart-meter readings recovered in five steps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the gain-tuned nonlinear PID update stays stable and converges quickly on the tested random-missingness cases, because the paper gives no proof of convergence or stability for the modified SGD rule and does not test the structured gaps that sensor failures actually produce.","fun_headline_variants_meta":{"raw":{"variants":["PID-tuned tensor model fills smart-meter gaps faster","Nonlinear PID speeds up missing-data tensor recovery","Five iterations to accurate NILM data recovery","PSO-tuned PID boosts tensor factorization convergence","Missing smart-meter readings recovered in five steps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000425,"raw_usage":{"total_tokens":2249,"prompt_tokens":1086,"completion_tokens":1163,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":702,"completion_tokens_details":{"reasoning_tokens":1092}},"tokens_in":702,"tokens_out":1163,"duration_ms":8912,"temperature":1.0,"reasoning_tokens":1092,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:08:22.210821+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run NPIL on UK-DALE at 15% density with entries missing in contiguous blocks, for example all three measurement channels absent for several consecutive hours to mimic sensor dropout, and compare against M3; if RMSE degrades to baseline levels or the training curve oscillates, the claim that the method recovers missing NILM data under sensor failure over-generalizes.","supporting_citations":[{"cited_title":"A nonlinear PID-incorporated adaptive stochastic gradient descent algorithm for latent factor analysis,","cited_arxiv_id":null,"evidence_quote":"Supplies the nonlinear PID-incorporated adaptive SGD rule that NPIL applies to tensor factorization."},{"cited_title":"Hierarchical particle swarm optimization-incorporated latent factor analysis for large-scale incomplete matrices,","cited_arxiv_id":null,"evidence_quote":"Provides the PSO-incorporated latent factor analysis scheme used to adapt NPIL's gain parameters."},{"cited_title":"Temporal pattern-aware QoS prediction via biased non-negative latent factorization of tensors,","cited_arxiv_id":null,"evidence_quote":"Defines M1, the biased non-negative LFT baseline whose RMSE and MAE NPIL must beat."},{"cited_title":"Multi-dimensional QoS prediction for service recommendations,","cited_arxiv_id":null,"evidence_quote":"Defines M2, the tensor-based QoS prediction baseline used in the runtime comparison."},{"cited_title":"A tensor-based approach for the QoS evaluation in service-oriented environments,","cited_arxiv_id":null,"evidence_quote":"Defines M3, the CP decomposition baseline reported as the strongest non-NPIL accuracy competitor."},{"cited_title":"Tensor factorization with sparse and graph regularization for fake news detection on social networks,","cited_arxiv_id":null,"evidence_quote":"Defines M4, the sparsity-and-graph-regularized tensor decomposition baseline."},{"cited_title":"Outlier-resilient web service QoS prediction,","cited_arxiv_id":null,"evidence_quote":"Defines M5, the Cauchy-loss outlier-resilient QoS prediction baseline."},{"cited_title":"It’s different: insights into home energy consumption in India,","cited_arxiv_id":null,"evidence_quote":"Supplies the iAWE dataset used for all D1 missing-data experiments."},{"cited_title":"The UK-DALE dataset, domestic appliance-level electricity demand and whole-house demand from five UK homes,","cited_arxiv_id":null,"evidence_quote":"Supplies the UK-DALE dataset used for all D2 missing-data experiments."}],"review_version":1}