{"id":"e18cb3d1-79d0-4b55-847a-9cc4739485d3","arxiv_id":"2607.22862","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A parameterized CEC 2017 benchmark with independent bias, shift, and rotation controls shows hMPA's results change most when shift and rotation act together, with no universal ordering of transformation effects.","lead":"This paper separates the three standard transformations of the CEC 2017 optimization benchmark and runs one hybrid optimizer through all eight combinations. It finds that shifting the optimum hurts results, while rotation alone does little harm, and that effects depend strongly on the test function and dimension.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Parameterization of composition functions is underspecified: Section 3 toggles a single o_i/M_i, but CEC 2017 F20–F30 have ten component-level shifts and rotations; if those are not zeroed/identity for C(000)/C(010)/C(001), isolation fails for 11 of 29 functions.","rationale":"The reader gave CONDITIONAL, and I agree. The most load-bearing risk is not the statistical machinery or the missing MPA baseline (both acknowledged by the authors) but the unverified implementation of the parameterization layer for composition functions. Section 3's notation is defined for basic functions and does not describe how the component-level shift/rotation data in F20–F30 are toggled. Without code or an explicit statement, a reader cannot tell whether the 'unshifted' configurations are truly unshifted for 11/29 functions. If this fails, the central methodological claim of independent control collapses for one third of the benchmark, and the observed rankings for composition functions could be artifacts. The concrete test above settles it directly. The reader's weakest_assumption already pointed at this; no change of verdict is needed beyond keeping CONDITIONAL pending the check.","tokens_in":29365,"tokens_out":10683,"duration_ms":98308,"concrete_test":"Obtain the authors' parameterized implementation and inspect cec17_param_func for composition functions F20–F30; verify that for C(000) the component shift matrix (D×10) is all zeros and all 10 rotation matrices are identity, and that for C(010) the original shift matrix is restored while rotations remain identity. As an empirical check, run C(000) on F21 and confirm the global optimum is at x=0 (or that evaluation at x=0 equals the function minimum). If component-level transforms remain active in C(000), the isolation claim fails for the composition class.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 defines the controlled form as F_i(x)=f_i(M_i(x-o_i))+F_i^* and states that 'the parameterization is limited to the controlled activation of F_i^*, o_i, and M_i'. For hybrid and composition functions the text says only that original variable partitioning, shuffle data, and component aggregation are preserved. In the CEC 2017 reference implementation [1], composition functions F20–F30 use N=10 components, each with its own shift vector (o is a D×N matrix) and rotation matrix (M is a D×D×N array). The paper never specifies whether setting 'shift to a zero vector' and 'rotation to identity' replaces all component-level transforms or only a single outer-level o_i/M_i. If the latter, then C(000), C(010), and C(001) would not remove shift/rotation for 11 of 29 functions, so the reported function-wise differences—especially the central claims that shift generally worsens outcomes and that shift+rotation differs most consistently from control—would be contaminated by unresolved component transforms for those functions. Since no code or raw data is released (Data Availability), the ambiguity cannot be resolved from the manuscript. This is the weakest link in the claim that the parameterization isolates bias, shift, and rotation independently.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces a parameterized variant of the CEC 2017 benchmark in which bias, shift, and rotation are controlled independently, producing eight configurations C(000)-C(111), while the original functions and transformation data are said to be preserved. Using this framework, the authors diagnose the hybrid Marine Predators Algorithm (hMPA), whose predicted-candidate operator depends on raw objective values and coordinate-wise reconstruction. DSC and eDSC are adapted from algorithm-level to configuration-level comparison, and the paper reports an exhaustive analysis of all eight configurations, all 56 three-configuration subsets, and control comparisons against C(000), across 29 functions at dimensions 10-100 with 30 runs under a budget of 10000*dim evaluations. The main findings are that DSC and eDSC detect statistically significant global differences at every dimension, that configuration rankings vary across functions and preclude a consistent ordering, that shift generally worsens objective-value outcomes, that isolated rotation and isolated bias have weak effects, and that shift-rotation combinations differ most consistently from the control. The authors explicitly hedge that the results are limited to hMPA, the chosen parameters, 30 runs, the fixed budget, and the analyzed functions, and they note that no factorial interactions or operator attribution are established.","tokens_in":29658,"tokens_out":23557,"duration_ms":203916,"significance":"If the reported results hold, the paper's main value is methodological: a parameterization layer that independently controls bias, shift, and rotation in CEC 2017, combined with a configuration-level adaptation of DSC/eDSC applied exhaustively (8 configurations, 56 three-way subsets, control comparisons, 29 functions, four dimensions). The exhaustive design, the explicit numerical handling of rank-deficient covariance matrices, and the repeated acknowledgment of the bias-offset confound in RankY, the informational limits of 30 solutions in 100 dimensions, and the absence of an MPA-only baseline are strengths; the paper does not overclaim factorial interactions and presents the conclusions as fixed-budget results. The empirical novelty is moderate, being a case study of one hybrid algorithm with fixed parameter settings, and the function-level reversals (e.g., C(010) ranking best on F23, F27, and F29) are the most interesting output. Overall significance is moderate and depends on the parameterization being correct for the hybrid and composition functions, which is exactly the point that the manuscript leaves unspecified.","major_comments":[{"comment":"The parameterization of hybrid and composition functions is not specified at the component level. The CEC 2017 reference implementation [1] stores component-level shift and rotation data for the hybrid functions F11-F20 and the composition functions F21-F30 (o and M are D x N and D x D x N arrays for N=10 components). Section 3 states only that \"the parameterization is limited to the controlled activation of F_i^*, o_i, and M_i\" and describes the deactivation as \"using a zero vector\" and \"using an identity matrix with a dimension consistent with the given function\"; this does not establish whether every component-level shift is zeroed and every component-level rotation is set to identity for configurations such as C(000), C(001), and C(010). If only top-level transformations were toggled, the isolation of shift and rotation would fail for 20 of the 29 test functions, which would directly contaminate the central claims that shift generally worsens objective-value outcomes, that isolated rotation causes no systematic deterioration, and the function-level reversals highlighted for F23, F27, and F29. Since no code is released, the ambiguity cannot be resolved from the manuscript; the authors should specify the component-level substitution explicitly and/or release the implementation for verification.","section":"Section 3 (Eq. (1)), Table 1"},{"comment":"The DSC control p-values for C(100), C(101), C(110), and C(111) versus C(000) are reported as identical (7.8E-08) at every dimension. Table 3 shows that the function-wise rank differences for these four configurations are not identical (for example, at dim=10 on F9 the differences from C(000) are 3.5, 5.5, 4.5, and 6.5, respectively), so the Wilcoxon statistics on RankY values would not generally coincide; the same holds across dimensions, where the rank patterns also differ. The authors should report the exact quantities paired (per-function RankY values or aggregated raw values), the test statistics, and the tie-handling procedure, and verify the printed p-values; as reported, these entries are not reproducible from the displayed ranks.","section":"Section 7.3, Table 7"}],"minor_comments":[{"comment":"Section 2 contains an unprocessed LaTeX command (\"textbf Centre-bias\"), and Section 9 contains corrupted text (\"comp etition\", \"we- re evaluated\", \"Analogical transfor - mation\"); the manuscript should be carefully proofread before resubmission.","section":"Sections 2 and 9"},{"comment":"The single figure is numbered \"Figure 8\" although no Figures 1-7 exist, so it should be renumbered; additionally, the convergence plots use raw objective values, and error-relative-to-optimum curves would be more informative, as the paper itself acknowledges in Section 8.","section":"Figure 8"},{"comment":"The combined significance value pvalue = 1 - prod(1-p_j) is applied to seven Wilcoxon comparisons that share the same control data, so the constituent p-values are not independent; the text should state that this combined value is a descriptive summary rather than a calibrated meta-test.","section":"Section 6"},{"comment":"The Data Availability statement says the datasets are available from the corresponding author upon reasonable request, but the paper's explicit claim of a reproducible diagnostic protocol would be better supported by releasing the parameterized benchmark layer and the adapted DSC/eDSC scripts, which would also resolve the component-level ambiguity raised in the first major comment.","section":"Data Availability"},{"comment":"The statement that reversals for composition functions \"do not contradict\" the shift mechanism would be easier to evaluate if the paper reported the magnitudes behind the rank reversals (for example, median final objective values per configuration for F23, F27, and F29), since DSC ranks alone conflate small and large differences.","section":"Section 8"},{"comment":"DSC comparisons involving bias-containing configurations (C(100), C(101), C(110), and C(111)) include the imposed additive offset, as the Discussion acknowledges; the caveat should also appear at the point of the first such comparison in the Results, and a DSC re-analysis on error-to-optimum values would strengthen the bias-related conclusions.","section":"Sections 7 and 8"},{"comment":"Verify that all listed references are cited in the body text, since reference [12] does not appear to be cited.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on the authors' prior work ([25] and [29]), and the parameterization concept is not entirely new relative to the CEC 2021 binary-toggle logic; the contribution lies in the exhaustive application and the diagnostic framing. The identical DSC p-values in Table 7 deserve a careful audit by the authors before publication, as they may reflect a computation or reporting artifact. No concerns about research integrity arise from the manuscript itself; the primary issues are technical specification, reproducibility, and presentation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, the paper is a careful empirical study of how hMPA responds to controlled bias, shift, and rotation in CEC 2017. It does what it says: introduces a parameterized benchmark layer, runs 29 functions at four dimensions with eight configurations, and applies DSC and eDSC to rank configurations. The central finding — that configurations differ statistically at every dimension but rankings vary by function so no consistent ordering exists — is supported by the reported tests. The paper earns credit for the exhaustive enumeration of all 56 three-configuration subsets, the control comparisons, and the explicit discussion of limitations: the additive bias offset in RankY, the rank-deficient 100-dimensional covariance from 30 runs, and the absence of an MPA-only baseline. These are acknowledged honestly.\n\nThe biggest soft spot is the parameterization of composition functions, F20–F30. In the CEC 2017 reference implementation those functions use ten component-level shift vectors and rotation matrices. Section 3 says the parameterization toggles F*_i, o_i, and M_i while preserving component aggregation, but it never states whether setting shift to zero and rotation to identity replaces all component-level transforms or only a single outer-level pair. If the latter, C(000), C(010), and C(001) would still contain active component shifts and rotations for 11 of 29 functions, which would contaminate the central claims about shift and rotation effects. The paper does not release code or raw data, so the ambiguity cannot be resolved externally. I think this is a genuine gap, but it is a missing specification rather than a demonstrated error; the authors may well have zeroed all component transforms.\n\nMinor points: the identical Wilcoxon p-values for several configurations in Table 7 are explained by identical signed-rank patterns, but it would help to show effect sizes. The 56 subset tests are not globally corrected, though the authors appropriately label them diagnostic. The convergence plots are consistent with, but not independent of, the statistical results.\n\nWho is this for: anyone doing benchmark parameterization or sensitivity analysis of metaheuristics, and people working with DSC and eDSC. It is a solid contribution that deserves a serious referee. I would recommend major revision or at least a mandatory clarification of the composition-function handling, ideally with code release.","headline":"Solid, well-scoped empirical diagnostic study of hMPA on parameterized CEC 2017; the main claims hold, but the composition-function parameterization is underspecified and needs clarification or code release.","tokens_in":30221,"tokens_out":3297,"would_cite":true,"duration_ms":28233,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hybrid optimizer's results depend on which benchmark transformations are active, with shift doing most of the damage and no consistent ordering across functions.","keywords":["CEC 2017 benchmark","benchmark parameterization","transformation sensitivity","metaheuristic optimization","Deep Statistical Comparison","extended Deep Statistical Comparison","hybrid Marine Predators Algorithm","bias-shift-rotation configurations"],"falsifier":"Re-run the standard CEC 2017 benchmark and verify that the full configuration C(111) in the parameterized implementation produces identical result distributions; then check that deactivating shift leaves the optimum at the search-space center for hybrid and composition functions, not only for simple functions. If either check fails, the shift and rotation attributions are artifacts of the parameterization layer rather than properties of hMPA.","tokens_in":29141,"feed_emoji":"📊","tokens_out":9169,"duration_ms":78779,"temperature":0.7,"pith_summary":"The paper sets out to explain why a hybrid optimizer's results change when the transformations built into a benchmark are switched on or off. It introduces a parameterized version of the CEC 2017 benchmark in which bias, shift, and rotation are controlled independently while the original functions and transformation data are preserved, and it applies this layer to the hybrid Marine Predators Algorithm (hMPA), whose predicted-candidate operator reconstructs solutions from numerical objective values and coordinate-wise models. Using Deep Statistical Comparison (DSC) and extended DSC (eDSC) re-purposed from algorithm ranking to configuration ranking, it claims that the eight bias-shift-rotation configurations differ statistically significantly at every dimension tested, yet the ordering is not consistent across the 29 functions. The concrete pattern is that shift generally worsens objective values, isolated rotation causes no systematic deterioration, isolated bias has little effect on final-solution distributions, and configurations combining shift with rotation differ most consistently from the untransformed control. This matters because the standard benchmark activates all three transformations at once, so ordinary rankings cannot say which part of a landscape an algorithm is actually responding to.","feed_headline":"Shift hurts hybrid optimizer; rotation alone doesn't","feed_subtitle":"A parameterized CEC 2017 benchmark shows the optimizer's rankings vary across functions and dimensions.","key_machinery":"The load-bearing object is the parameterization layer: each benchmark function is written as $F_i(x)=f_i(M_i(x-o_i))+F^*_i$, with a binary configuration vector $C=(C_1,C_2,C_3)$ deciding whether the bias $F^*_i$, the shift $o_i$, and the rotation matrix $M_i$ are active, so $C(000)$ is the untransformed control and $C(111)$ reproduces the standard CEC 2017 benchmark. The second piece is the re-purposed statistical machinery: DSC compares distributions of final objective values using pairwise distribution tests with a multiple-comparison correction, while eDSC compares final solution vectors using a multivariate E-test and covariance hypervolumes, with explicit numerical handling of rank-deficient covariance matrices when the dimension exceeds the number of runs. The mechanism under diagnosis is hMPA's predicted-candidate operator, which builds a one-dimensional linear regression of objective value on each coordinate and inverts it to reconstruct a candidate from a target value; that is why bias, shift, and rotation can probe different parts of the algorithm.","core_discovery":"The discovery is that hMPA is transformation-sensitive in a way that is statistically clear but not monotonically ordered. Across 29 CEC 2017 functions, four dimensions (10, 30, 50, 100), 30 runs per setting, and a fixed budget of 10000*dim evaluations, rank-based significance tests on both DSC and eDSC rankings show significant differences among the eight configurations at every dimension, and the same holds for nearly all 56 three-configuration subsets. The directional effects are separable: shift is the transformation that generally worsens final objective values; isolated rotation does not systematically degrade results; isolated bias barely changes the distribution of final solutions; and configurations containing both shift and rotation differ from the control most consistently. Crucially, the rankings flip from function to function, so the configurations cannot be arranged on a single difficulty scale, and the effect of a combined configuration is not the sum of its components' effects. The paper therefore claims the eight configurations are diagnostic conditions rather than ordered difficulty levels.","pith_inferences":["The predicted-candidate operator is the most plausible coupling point for the observed sensitivity: coordinate-wise inverse regression should be disturbed most by rotation, which mixes coordinates, and by shift, which changes the population-mean target values; a matched MPA-only ablation would likely show weaker transformation effects, though the paper runs no such baseline.","The weaker eDSC separation at dimension 100 may be an artifact of estimating a 100-dimensional distribution from only 30 final vectors, so the apparent convergence of configurations at high dimension should not be read as genuine similarity without more runs or projected summaries.","Subtracting each configuration's optimum from its final solutions before running eDSC would separate the deterministic displacement caused by shift from genuine distributional change; the paper itself signals this as a natural complement.","The configuration set is naturally a factorial design (bias by shift by rotation), so a factorial analysis of variance on ranks could estimate interactions formally; the paper declines to fit such a model, leaving interaction claims untested."],"forward_implications":["A standard CEC 2017 ranking cannot attribute performance to bias, shift, or rotation; configuration-level testing is required to know which transformation caused an outcome.","Benchmark designers using hMPA should expect shift, and especially shift combined with rotation, to be the conditions that separate behavior, while isolated rotation and bias are weak levers.","Because rankings vary across functions and dimensions, statements about a configuration being harder or easier must be qualified by the specific function and dimension.","The same parameterization and DSC/eDSC workflow can be transferred to other continuous optimizers and to CEC 2024 and CEC 2025 suites that reuse CEC 2017 problems.","Fixed-budget results are not asymptotic: configurations at different convergence stages when the budget ends may reorder under a larger budget, so transformation effects should be reported together with the budget."],"supporting_citations":[{"why":"Supplies the CEC 2017 function definitions, shift vectors, rotation matrices, and bias values that the parameterization layer preserves and toggles.","marker":"[1]"},{"why":"Defines DSC and eDSC, the distribution-comparison methods that the paper adapts from algorithm-level to configuration-level diagnosis.","marker":"[7]"},{"why":"Provides the base Marine Predators Algorithm implementation (phases, FADs effect, parameters) on which hMPA is built.","marker":"[8]"},{"why":"Introduces the binary bias-shift-rotation transformation logic for CEC 2021 that the parameterization layer follows.","marker":"[18]"},{"why":"Documents the CEC 2021 protocol's treatment of shift-free scenarios, supporting the diagnostic use of configurations without shift.","marker":"[20]"},{"why":"Defines the predicted-candidate operator and the hMPA variant central to the sensitivity study, including the newNopar setting.","marker":"[25]"},{"why":"Shows that objective-function transformations differentiate hybrid metaheuristics, motivating the hMPA case study.","marker":"[29]"}],"fun_headline_variants":["Shift degrades hMPA; rotation alone doesn't","Benchmark tweaks reveal hMPA's variable sensitivity","hMPA's response to CEC shifts: not a simple scale","Isolated rotation doesn't hurt hMPA; shift does","No single difficulty order for hMPA under CEC transforms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The diagnosis assumes that turning off a transformation only removes that transformation (zero shift vectors and identity rotations must not alter the hybrid or composition functions in unintended ways) and that the hMPA implementation faithfully combines the published base MPA and predicted-candidate operator, with no MPA-only baseline to rule out inherited base-algorithm sensitivity.","fun_headline_variants_meta":{"raw":{"variants":["Shift degrades hMPA; rotation alone doesn't","Benchmark tweaks reveal hMPA's variable sensitivity","hMPA's response to CEC shifts: not a simple scale","Isolated rotation doesn't hurt hMPA; shift does","No single difficulty order for hMPA under CEC transforms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000392,"raw_usage":{"total_tokens":2104,"prompt_tokens":1035,"completion_tokens":1069,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":651,"completion_tokens_details":{"reasoning_tokens":983}},"tokens_in":651,"tokens_out":1069,"duration_ms":7117,"temperature":1.0,"reasoning_tokens":983,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:28:05.855841+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the standard CEC 2017 benchmark and verify that the full configuration C(111) in the parameterized implementation produces identical result distributions; then check that deactivating shift leaves the optimum at the search-space center for hybrid and composition functions, not only for simple functions. If either check fails, the shift and rotation attributions are artifacts of the parameterization layer rather than properties of hMPA.","supporting_citations":[{"cited_title":"Awad, M.Z","cited_arxiv_id":null,"evidence_quote":"Supplies the CEC 2017 function definitions, shift vectors, rotation matrices, and bias values that the parameterization layer preserves and toggles."},{"cited_title":"Eftimov and P","cited_arxiv_id":null,"evidence_quote":"Defines DSC and eDSC, the distribution-comparison methods that the paper adapts from algorithm-level to configuration-level diagnosis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the binary bias-shift-rotation transformation logic for CEC 2021 that the parameterization layer follows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the CEC 2021 protocol's treatment of shift-free scenarios, supporting the diagnostic use of configurations without shift."},{"cited_title":"Oszust, G","cited_arxiv_id":null,"evidence_quote":"Defines the predicted-candidate operator and the hMPA variant central to the sensitivity study, including the newNopar setting."},{"cited_title":"Sroka and S.T","cited_arxiv_id":null,"evidence_quote":"Shows that objective-function transformations differentiate hybrid metaheuristics, motivating the hMPA case study."}],"review_version":1}