{"id":"37c8a432-97e3-411b-ae6b-4debac3a5fb2","arxiv_id":"2607.09318","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"EcoKube provides a configurable event-driven simulator and reference policy that cuts estimated emissions ~45% versus default Kubernetes under synthetic hybrid edge-cloud scenarios, with modest makespan cost.","lead":"EcoKube is a deterministic simulator for comparing carbon-aware Kubernetes-style schedulers on heterogeneous edge-cloud sites with CI, PUE, and hardware fit. It gives operators a reproducible way to test sustainability policies before real deployment.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The 45% emissions cut is largely an in-model optimization of the identical CI/PUE signals used in scoring, so it does not yet evidence real pre-deployment ranking power.","rationale":"The paper’s architectural contribution (deterministic event-driven simulator, Kubernetes-style feasibility + pluggable Score hooks, open seeds/configs) is sound and useful for controlled comparison. The load-bearing soft spot is exactly the one the reader flagged: the emissions quantity used both for ranking and for the reported outcome is model-internal, synthetic, and uncalibrated. No internal inconsistency or calculation error is evident; the 45% figure is simply not yet evidence of transferable sustainability ranking. Because the authors already list real-cluster calibration and broader baselines as future work, and the contribution is framed as a testing workflow rather than a new carbon-physics result, the reader’s CONDITIONAL verdict remains appropriate. No stronger objection (e.g., non-determinism, missing feasibility filtering, or fabricated baselines) appears in the text.","tokens_in":10730,"tokens_out":583,"duration_ms":28714,"concrete_test":"Using the public GitHub artifacts, re-run the exact scenario keys of Table 4 after replacing the policy’s energy estimator Ê with an independent utilization-to-power model (e.g., published SPECpower or RAPL-derived curves for the CPU/GPU classes) that is never supplied to any scorer; if ecokube-pol’s relative emissions reduction falls below ~15% or the ordering among baselines changes, the headline delta is an artifact of closed-loop optimization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest experimental claim (45.15% drop in estimated emissions, Table 4: 32.14 kg → 17.63 kg) treats total_ci_cost_g (energy × PUE × time-aligned CI, optionally scaled by k_s) as an external outcome (§4.1). Yet EcoKube Policy’s node score (Eq. 3) and site ranking directly embed the same product Ê·PUE·CI with non-trivial weight (w_C = 0.21 plus site-level terms). Workloads are synthetic, energy estimates uncalibrated, and the topology is a single fixed 3-site/8-node substrate. Consequently the large gap versus k8s/KEIDS/TOPSIS primarily shows that the reference policy optimizes the simulator’s own objective more aggressively than the baselines under matched seeds; it does not yet demonstrate that the framework produces rankings that transfer to real multi-site carbon impact. This undercuts the “reproducible way to compare \tau before deployment” claim until the metric is externally validated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"EcoKube is a Go-based discrete-event simulator for comparing sustainability-aware scheduling policies on heterogeneous federated edge–cloud topologies. It models site-level CI/PUE (and a normalisation factor k_s), node-level feasibility (resources, accelerators, latency), and exposes Kubernetes-style hard filtering plus a pluggable Score hook. A reference two-stage weighted policy (site then node; Eqs. 3–4) is compared under matched seeds to default Kubernetes, KEIDS, and TOPSIS/KCSS on synthetic batch workloads over a fixed three-site (NL/FR/DE), eight-node substrate. In the reported slice, EcoKube Policy cuts the model’s operational emissions estimate by ~45% relative to k8s (32.14 kg → 17.63 kg mean per ~900-job run; Table 4) with ~1.35% makespan increase. The stated contribution is architectural and experimental: a reproducible pre-deployment comparison workflow, not a new optimisation paradigm.","tokens_in":11112,"tokens_out":1272,"duration_ms":25441,"significance":"If the framework is sound and usable, it fills a real gap: carbon-aware placement work often lacks a shared, Kubernetes-compatible test harness that jointly models site CI/PUE and node heterogeneity. Strengths include deterministic matched evaluation, public artifacts (GitHub), explicit policy hooks, and honest framing of EcoKube Policy as a reference instantiation. The result is primarily an engineering/methodology contribution for the sustainable systems and edge–cloud scheduling community; transfer of the 45% emissions figure to real multi-site carbon impact is not yet established and should not be the main claim.","major_comments":[{"comment":"§4.1–4.2 and Table 4: The headline 45.15% emissions reduction treats total_ci_cost_g (energy × PUE × time-aligned CI, cf. Eq. 1) as an external outcome, yet EcoKube Policy’s site and node scores (Eqs. 3–4) directly embed the same Ê·PUE·CI product with non-trivial weight (w_C = 0.21 plus site-level terms). Under synthetic energy estimates and a single fixed topology, the large gap vs k8s/KEIDS/TOPSIS mainly shows more aggressive optimisation of the simulator’s own objective. For the central claim of a “reproducible way to compare … before deployment,” reframe results as in-model ranking behaviour, report sensitivity when the evaluation metric is partially decoupled from scoring inputs, or add an external/held-out carbon proxy.","section":null},{"comment":"§3.1.3 Eq. (3) and §4.1: Estimated IT energy Ê_{w,n} is load-bearing for both ranking and the reported emissions metric, but the manuscript does not specify how energy is computed (power model, utilisation, GPU vs CPU, contention, idle power, or duration). Without this, absolute kg figures and cross-policy deltas in Table 4 cannot be audited or reproduced from the text alone. Document the energy model, its parameters, and any calibration assumptions; if purely synthetic, state that explicitly and bound sensitivity.","section":null},{"comment":"§4.1 and Table 3: Evaluation uses one three-site/eight-node topology, three synthetic mix presets, and no real-cluster power or placement traces. §4.3 acknowledges this, but the abstract and conclusions still present the workflow as ready for pre-deployment comparison. Either expand to at least one additional topology/workload family (or a small empirical replay) or narrow the claim to “controlled synthetic comparison under hybrid heterogeneity,” with quantitative limits on generalisation.","section":null},{"comment":"Table 2 / §4.1 baselines: KEIDS and TOPSIS/KCSS are reimplemented inside EcoKube. Faithfulness of those ports (objective terms, interference model for KEIDS, criteria set for TOPSIS) is not validated against original papers or reference code. Mis-implementation would inflate EcoKube Policy’s relative gain. Provide implementation notes, parameter mapping, and a short sanity check that baseline behaviour matches published intent under a simple scenario.","section":null}],"minor_comments":[{"comment":"§3.1.4 vs Table 3: Site-level weights are written as α, β, γ in the table but as a weight vector W and separate site parameters in the text; align notation throughout.","section":null},{"comment":"Figure 2: Axis labels repeat the metric name; add error bars or IQR over the 50 repetitions so variance is visible.","section":null},{"comment":"Eq. (1): k_s is introduced as compensating for “measurement and attribution differences” but never given numerical values or a selection method in the evaluation setup.","section":null},{"comment":"Abstract and §1: “TOPSIS/KCSS” is slightly ambiguous; Table 2 lists topsis with citation [8] (KCSS)—clarify whether one or two baselines are used.","section":null},{"comment":"§4.1: Warm-up of 30 minutes is stated but not justified relative to arrival rates and job counts; a one-sentence rationale would help.","section":null},{"comment":"References: several arXiv and survey items are fine for a workshop, but ensure KEIDS venue/year consistency with the IEEE IoT Journal citation as printed.","section":null}],"recommendation":"major_revision","confidential_remarks":"Workshop-length systems paper; contribution is a useful harness if claims stay architectural. The 45% number is easy to over-read in publicity. Scope fit for TDIS is reasonable. I would accept after the metric framing, energy-model disclosure, and baseline-fidelity notes are fixed—even without a full real-cluster study—if the abstract/conclusions are toned to match the evidence."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The paper’s real product is EcoKube: a from-scratch deterministic discrete-event simulator with Kubernetes-style feasibility filtering, pluggable Score hooks, site/node adapters for CI/PUE/topology, and open artifacts. That is the new piece. Carbon-aware placement, KEIDS, TOPSIS/KCSS, and CI×PUE models are all prior art they cite correctly; the contribution is the matched, seed-controlled harness for heterogeneous edge–cloud topologies, not a new scheduling theory.\n\nWhat they do well is engineering hygiene. Same inputs and arrivals for every policy, fixed seeds, 50 reps, clear absolute table (Table 4), weight sweeps treated as config knobs rather than claimed optima, and an explicit statement that EcoKube Policy is a reference instantiation, not a new paradigm. The two-stage site-then-node scoring (Eqs. 3–4) is transparent and easy to re-implement. Citation pattern is solid and the math is just weighted sums—no load-bearing errors.\n\nThe soft spot is exactly the stress-test point, and it is real but proportionate. The 45% drop (32.14 kg → 17.63 kg) is total_ci_cost_g recomputed from the same energy×PUE×CI signals the policy optimizes (w_C = 0.21 plus site terms). Workloads are synthetic, energy uncalibrated, topology fixed at three sites/eight nodes. So the gap mainly shows that their weighted policy chases the simulator’s own objective harder than k8s/KEIDS/TOPSIS under matched seeds. Authors already flag this and list real-cluster calibration as future work; it does not sink the framework claim, but it does mean the “before deployment” ranking power is still unproven.\n\nThis is for people building or evaluating carbon-aware Kubernetes extensions who need a controlled sandbox. Not for carbon-physics or large-scale systems results. I would send it to peer review: the artifact is concrete, the evaluation is honest about its limits, and the community needs exactly this kind of reproducible harness. Engage if you care about the tooling; treat the 45% figure as an internal demonstration, not transfer evidence.","headline":"Useful open simulator for carbon-aware K8s-style policy comparison; the 45% emissions number is mostly in-model optimization, not yet external proof.","tokens_in":11699,"tokens_out":537,"would_cite":true,"duration_ms":6289,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"EcoKube is a configurable, Kubernetes-compatible simulator that makes carbon-aware scheduling policies reproducible and comparable on heterogeneous edge–cloud topologies before real deployment.","keywords":["sustainable computing","carbon-aware computing","federated computing systems","Kubernetes scheduling","multi-objective optimisation","Carbon Intensity","discrete-event simulation","edge-cloud"],"falsifier":"Deploy the identical EcoKube Policy and the three baselines on a real multi-site Kubernetes federation equipped with calibrated power meters and the same carbon-intensity traces; if measured carbon savings collapse or reverse while the simulator still reports a large drop, the claim that the framework correctly ranks real sustainability impact is falsified.","tokens_in":11572,"feed_emoji":"🌱","tokens_out":919,"duration_ms":20162,"temperature":0.7,"pith_summary":"Hybrid edge–cloud systems burn different amounts of carbon depending on where and when a job runs, because grid carbon intensity, facility efficiency, and hardware (including GPUs) all vary. Existing carbon-aware schedulers lack a shared, controlled way to measure those trade-offs under the same mixed topology. EcoKube supplies that missing testbed: a deterministic event-driven simulator with site-level sustainability signals, node-level feasibility filters, and pluggable policy hooks that stay close to Kubernetes practice. Its reference policy, which first picks cleaner sites then ranks nodes by energy, emissions, latency and hardware fit, cuts the reported operational emissions estimate by roughly 45 percent versus default Kubernetes on synthetic batch workloads while only slightly lengthening makespan. The paper’s claim is architectural and experimental: give operators a reproducible workflow so they can compare sustainability-aware policies before they risk them in production.","feed_headline":"Simulator cuts estimated carbon 45% vs Kubernetes defaults","feed_subtitle":"EcoKube lets teams test carbon-aware placement on mixed edge-cloud hardware before they ship it","key_machinery":"The two-stage EcoKube scoring rule: first rank candidate sites by normalised carbon intensity, PUE, site factor and optional network penalty; then, inside the chosen site, rank feasible nodes by a weighted sum of estimated IT energy, emissions (energy × PUE × CI), latency penalty and accelerator-fit term (Eqs. 3–4), after Kubernetes-style hard-constraint filtering.","core_discovery":"A modular discrete-event simulator that models both site-level carbon and efficiency signals and node-level heterogeneity can turn sustainability-aware scheduling into a controlled, Kubernetes-compatible experiment. Under the reported three-site synthetic scenarios, the framework’s reference weighted-sum policy reduces the mean operational emissions estimate from 32.14 kg to 17.63 kg per roughly 900-job run (a 45.15 percent drop relative to default Kubernetes) while increasing makespan by only 1.35 percent.","pith_inferences":["The same site/node split could be lifted into a live multi-cluster Kubernetes scheduler once power and CI adapters are calibrated against measured traces.","Because every run is deterministic and seed-controlled, the simulator can serve as a regression harness for any new multi-objective ranking method that claims carbon awareness.","Adding delay-tolerant temporal shifting inside the same topology would let researchers quantify the extra gain from combining spatial and temporal levers under identical conditions."],"forward_implications":["Operators can rank carbon-aware heuristics under controlled edge–cloud heterogeneity before writing Kubernetes extensions.","Explicit site-level CI/PUE and node-level accelerator modelling changes observed carbon–performance trade-offs relative to homogeneous-cluster assumptions.","The same seed-controlled workflow supports weight, arrival-rate and workload-mix sweeps that stress-test policy rankings.","Substantial reductions in the reported emissions estimate are achievable without large makespan penalties when both site and node signals are used."],"fun_headline_variants":["EcoKube trims estimated carbon 45% vs Kubernetes defaults","Carbon-aware reference policy cuts emissions 45% in EcoKube sims","EcoKube simulator: 45% lower ops carbon than default K8s scheduler","Heterogeneous edge-cloud sim shows 45% carbon drop vs Kubernetes","Weighted-sum policy in EcoKube halves carbon estimate vs default K8s"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The simulator’s own location-aware emissions number (energy times site PUE times time-aligned carbon intensity) is treated as a valid external measure of real sustainability impact even though the workloads are synthetic, the topology is fixed, and no calibrated power measurements from live multi-site clusters are used.","fun_headline_variants_meta":{"raw":{"variants":["EcoKube trims estimated carbon 45% vs Kubernetes defaults","Carbon-aware reference policy cuts emissions 45% in EcoKube sims","EcoKube simulator: 45% lower ops carbon than default K8s scheduler","Heterogeneous edge-cloud sim shows 45% carbon drop vs Kubernetes","Weighted-sum policy in EcoKube halves carbon estimate vs default K8s"]},"model":"grok-4.5","effort":"low","cost_usd":0.0055,"raw_usage":{"total_tokens":1492,"prompt_tokens":767,"num_sources_used":0,"completion_tokens":99,"cost_in_usd_ticks":55000000,"prompt_tokens_details":{"text_tokens":767,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":626,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":767,"tokens_out":99,"duration_ms":7644,"temperature":1.0,"reasoning_tokens":626,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T03:54:16.258582+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Deploy the identical EcoKube Policy and the three baselines on a real multi-site Kubernetes federation equipped with calibrated power meters and the same carbon-intensity traces; if measured carbon savings collapse or reverse while the simulator still reports a large drop, the claim that the framework correctly ranks real sustainability impact is falsified.","supporting_citations":[],"review_version":1}