{"id":"6407a372-06e3-4259-abf5-443fafac6746","arxiv_id":"2508.07691","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A first empirical study of the energy consumption of surrogate-assisted particle swarm optimization, proposing energy and surrogate accuracy as evaluation axes.","lead":"This paper measures the processor and memory energy use of several particle-swarm optimization variants that rely on neural-network surrogates, and compares that with how accurately the surrogates guide the search. It is a first look at whether energy consumption should join accuracy and runtime in judging optimization algorithms.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Energy-measurement reliability is the unverified linchpin; with the supplied corrupt text, the comparative energy claims cannot be checked.","rationale":"The reader's verdict is UNVERDICTED because the full text is corrupted. My stress-test identifies the same load-bearing assumption: the energy measurement setup must produce stable, comparable, and generalizable numbers. The paper's abstract makes a modest empirical claim — a first step toward energy-aware surrogate-assisted optimization — so the central contribution is entirely dependent on the validity and reproducibility of the energy measurements. Without access to the methods, we cannot check whether CPU/memory energy was properly separated, how many runs were used, or whether the reported differences exceed measurement noise. The 'acceptable solution' yardstick for surrogate accuracy is also vague; unless defined in the body, the energy/quality trade-off is uninterpretable. These concerns are not demonstrations of error; they are epistemic gaps. The paper may be perfectly sound once the readable text is available. Therefore, the correct verdict remains unchanged: UNVERDICTED, pending a proper reading. I agree with the reader's identification of the weakest assumption and recommend no change to the verdict based on this stress test.","tokens_in":14142,"tokens_out":3012,"duration_ms":35197,"concrete_test":"Obtain a readable copy of the paper and inspect the measurement section. Verify: (1) number of independent runs per configuration; (2) whether CPU and memory energy are obtained via hardware counters (e.g., RAPL/Perf) or external metering, and how memory energy is separated; (3) whether standard deviations or CIs are reported for energy values. If these are absent, perform a minimal replication: pick one benchmark and one PSO variant, measure CPU energy with Intel RAPL over 30 runs, and determine whether the reported ordering of surrogate variants is stable and significantly different from run-to-run noise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that surrogate-assisted PSO variants (pre-trained vs. retrained neural networks) exhibit measurably different processor and memory energy profiles while maintaining acceptable solution quality. This is an empirical, measurement-driven claim. For it to hold, the energy instrumentation must be valid: e.g., using hardware counters (RAPL or similar) or an external power meter, sampling at sufficient rate, isolating CPU vs. memory energy, and reporting enough repeated runs to distinguish signal from thermal and OS noise. The provided 'full text' is mojibake; we cannot confirm any of these details. The abstract offers no effect sizes, CIs, or run counts. Additionally, the 'acceptable solution' criterion for surrogate accuracy appears undefined, so the energy/accuracy trade-off cannot be interpreted. This is the weakest load-bearing point: if the energy measurements are noisy or hardware-specific, the entire comparison collapses; if they are solid, the paper delivers its modest 'first step.' No internal inconsistency is visible, but the evidence is currently inaccessible.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a first-step empirical study of the energy consumption of surrogate-assisted particle swarm optimization. It compares PSO variants using pre-trained and retrained neural-network surrogates, measuring processor and memory energy separately, and it assesses surrogate accuracy with respect to an 'acceptable solution' criterion. The authors argue that energy and surrogate accuracy should be considered alongside runtime and numerical quality. The abstract is readable and motivates the study, but the body of the paper is presented as unreadable mojibake in the provided submission, so the experimental setup, tables, and results could not be verified.","tokens_in":14365,"tokens_out":3195,"duration_ms":39969,"significance":"If the measurements are sound, this is a useful and timely contribution: energy profiling of surrogate-assisted metaheuristics, especially the separate treatment of processor and memory energy, is genuinely understudied. The proposed multidimensional assessment (quality, runtime, energy, surrogate accuracy) is a reasonable step for the community. However, the contribution is empirical, and no code, reproducible scripts, or machine-checked proofs are visible; the evaluation rests entirely on the experimental section, which is currently inaccessible.","major_comments":[{"comment":"The body of the paper is presented as mojibake; Sections containing the experimental setup, algorithms, tables, and results are not legible. The central claim is an empirical comparison of energy and accuracy, so the measurement instrumentation, benchmark definitions, surrogate training protocol, repetition counts, and result tables cannot be checked. This is load-bearing and must be corrected before the manuscript can be evaluated.","section":"Full text (as rendered)"},{"comment":"Even from the legible parts, no experimental reporting details are available: hardware platform, energy measurement method (RAPL, external power meter, OS counters, or other), sampling rate, number of runs, variance or confidence intervals, or statistical testing. Without these, the claimed differences between PSO versions are not supported, and the paper currently reads as a qualitative research agenda rather than a measurement study.","section":"Abstract and visible table fragments"},{"comment":"The yardstick for surrogate accuracy is undefined. What counts as an 'acceptable solution' (absolute error threshold, success rate, fixed budget, or something else) is essential for interpreting the energy-versus-accuracy trade-off. The manuscript should define this criterion explicitly per benchmark and report how many runs meet it.","section":"Abstract ('acceptable solution')"},{"comment":"The manuscript does not state whether surrogate accuracy is evaluated on training data, validation data, or held-out test data, nor is the retraining schedule described. If accuracy is measured on the data used to train the surrogate, the comparison is uninterpretable. This distinction is central to the claimed 'surrogate accuracy to properly drive the search' and must be clarified.","section":"Surrogate training/evaluation protocol"}],"minor_comments":[{"comment":"The phrase 'shed new light' is vague. The abstract should state at least one concrete quantitative finding, such as an observed energy difference or an accuracy level, to give readers a basis for judging the contribution.","section":"Abstract"},{"comment":"The visible table fragments lack labels, units, and sample sizes. Tables should include units (joules, watts), standard deviations or confidence intervals, and the number of repetitions per configuration.","section":"All tables"},{"comment":"The claimed 'first step toward a methodology' would be better supported by an explicit description of the proposed methodology (variables measured, normalization, reporting format) rather than a general appeal to holistic assessment.","section":"Introduction/Conclusion"},{"comment":"The submitted full text is unreadable due to encoding corruption. Authors should verify the PDF/TeX encoding before resubmission, as this currently prevents any substantive review of the experiments.","section":"Formatting"}],"recommendation":"major_revision","confidential_remarks":"The unreadable body may be an artifact of the submission rendering, but as provided it blocks verification. I recommend asking the authors to resubmit a legible version and to add the experimental reporting details (hardware, instrumentation, repetitions, statistics, definition of acceptable solution). If the resubmission does not include these, the paper would be difficult to accept as an empirical study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The abstract makes a real, understudied point: energy profiling of surrogate-assisted PSO, separating processor from memory costs, and tying it to surrogate accuracy. That is a defensible first step, and the writing is honest and scoped. The authors do not overclaim; they say \"first step\" and mean it.\n\nThe problem is that the full text I was given is corrupted beyond recovery. So I cannot check the measurement setup, the numbers, or the statistical treatment. The stress-test note is right that if the energy figures are noisy or hardware-specific, the central comparison collapses. But that is an open question, not a demonstrated flaw. The abstract alone gives no effect sizes, no run counts, no hardware description, and no definition of what counts as an \"acceptable solution\"—which you would need to interpret the energy/accuracy trade-off.\n\nWhat the paper does well, as far as I can tell, is to frame a modest empirical contribution and position it as a step toward a broader methodology. The novelty claim is plausible: energy is rarely reported in surrogate-assisted metaheuristics, and separating processor from memory energy is a useful detail.\n\nThe soft spot is entirely in the evidence layer. I cannot verify whether the energy measurement is reliable, whether the benchmarks are representative, or whether the reported differences are real as opposed to noise. If the authors provide a readable PDF with details on hardware and methodology, the paper deserves a serious referee. Without that, it is just an abstract with a promising topic.\n\nSo: send it to review if you get a clean copy; otherwise it is unassessable. I would not cite it yet, but I would keep it on the list to check once the actual text is available.","headline":"A plausible first step on energy profiling for surrogate-assisted PSO, but the supplied text is unreadable—so the evidence can't be checked, and that's the whole ballgame.","tokens_in":14829,"tokens_out":2245,"would_cite":false,"duration_ms":25045,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper tries to establish that surrogate-assisted optimization should be judged by energy and surrogate accuracy, not just runtime and solution quality, and reports a first measurement study of particle swarm optimization (PSO) versions","keywords":["surrogate-assisted optimization","particle swarm optimization","neural network surrogates","energy profiling","processor and memory energy","surrogate accuracy","metaheuristics"],"falsifier":"Repeat the same benchmark suite and PSO variants on a different processor and memory configuration, keeping everything else fixed, and check whether the ordering of variants by processor and memory energy remains the same; if the ranking flips, the central energy comparison is not generalizable across hardware.","tokens_in":14053,"feed_emoji":"⚡","tokens_out":2235,"duration_ms":27170,"temperature":0.7,"pith_summary":"The paper argues that energy profile and surrogate accuracy are understudied but important dimensions for evaluating surrogate-assisted metaheuristics. It reports a first measurement study comparing particle swarm optimization variants, including versions with pre-trained and retrained neural-network surrogates, tracking both processor and memory energy. The authors position this as a first step toward a methodology that characterizes such algorithms by energy and surrogate accuracy in addition to time and numerical efficiency. If the paper is right, energy should become a standard reported metric when comparing optimization algorithms, and surrogate accuracy should be evaluated by its ability to drive search toward acceptable solutions, not just fit the training data.","feed_headline":"Surrogate PSO variants differ measurably in energy use","feed_subtitle":"First analysis tracks processor and memory energy plus surrogate accuracy across pre-trained and retrained neural-network PSO.","key_machinery":"The central objects are PSO variants: standard PSO, PSO with a pre-trained neural network as a fitness surrogate, and PSO with a retrained or continuously updated neural network surrogate. The evaluation machinery is the joint measurement of processor energy, memory energy, and surrogate accuracy, used to compare the variants' search behavior and output quality.","core_discovery":"On its own terms, the paper claims that surrogate-assisted PSO versions can be measurably distinguished by their energy consumption in the processor and memory, and that the accuracy of the neural-network surrogate meaningfully affects the search's ability to reach an acceptable solution. The authors treat these measurements as shedding new light on surrogate-assisted search and as the beginning of a more holistic assessment framework for optimization and learning techniques.","pith_inferences":["The energy cost of training the surrogate model itself is likely separate from the energy of using it during search; a full accounting that amortizes training energy could change the ranking between pre-trained and retrained variants.","The findings suggest a broader hypothesis: that energy profiles are algorithm-specific and can be used as a selection criterion for green or energy-constrained computing environments.","A natural extension would be to test whether the energy-accuracy ordering observed here generalizes across different hardware, operating systems, or problem classes; if not, the methodology still stands but the specific rankings may be hardware-dependent.","The 'acceptable solution' yardstick could be standardized (e.g., a fixed threshold relative to known optima) so that surrogate accuracy is comparable across studies."],"forward_implications":["Energy consumption should be reported alongside runtime and solution quality when comparing optimization algorithms.","Pre-trained and retrained surrogate models may exhibit different energy-accuracy tradeoffs, so the choice of surrogate training strategy becomes an energy-relevant design decision.","The proposed measurement approach can be applied to other surrogate-assisted metaheuristics, not just PSO.","Surrogate accuracy, measured by the ability to guide search to acceptable solutions, becomes a more meaningful quality metric than raw prediction error.","A first-step methodology emerges for holistic characterization of optimization techniques, covering time, numerical efficiency, energy, and surrogate accuracy."],"supporting_citations":[],"fun_headline_variants":["Surrogate PSO energy varies by network retraining","Energy profile separates surrogate PSO variants","First look at surrogate search energy and accuracy","PSO surrogates: energy differs, accuracy matters","Measuring energy and accuracy of surrogate PSO"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The energy measurements are stable and general enough that the reported differences between PSO versions come from the algorithms themselves, not from the specific hardware, operating system, or measurement setup used in the experiments.","fun_headline_variants_meta":{"raw":{"variants":["Surrogate PSO energy varies by network retraining","Energy profile separates surrogate PSO variants","First look at surrogate search energy and accuracy","PSO surrogates: energy differs, accuracy matters","Measuring energy and accuracy of surrogate PSO"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000145,"raw_usage":{"total_tokens":979,"prompt_tokens":674,"completion_tokens":305,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":418,"completion_tokens_details":{"reasoning_tokens":235}},"tokens_in":418,"tokens_out":305,"duration_ms":3868,"temperature":1.0,"reasoning_tokens":235,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:55:49.352913+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the same benchmark suite and PSO variants on a different processor and memory configuration, keeping everything else fixed, and check whether the ordering of variants by processor and memory energy remains the same; if the ranking flips, the central energy comparison is not generalizable across hardware.","supporting_citations":[],"review_version":1}