{"id":"a2218727-3c9c-4f20-a957-335439a28e83","arxiv_id":"2607.28407","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"A three-layer taxonomy (compute, network, application) plus novel continuum-specific metrics and acquisition requirements for evaluating distributed computing continuum systems.","lead":"This paper catalogs performance metrics for distributed computing continuum systems spanning cloud, edge, and IoT, with formulas and collection guidance. It gives practitioners a shared checklist so heterogeneous continuum designs can be compared more consistently.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-stated condition on unvalidated novel indices.","rationale":"The paper's strongest claim is that the taxonomy + formulations + acquisition requirements enable more transparent cross-layer DCCS evaluation. That claim is definitional; it holds if the organization is coherent and the formulas are usable, which they largely are. The reader's CONDITIONAL verdict already isolates the right soft spot—the unvalidated novel composites—and correctly keeps correctness risk low. My stress pass finds no deeper fracture (no equation that silently assumes homogeneity the model denies, no acquisition tag that contradicts the metric definition, no circular empirical claim). Therefore the verdict should remain CONDITIONAL with no further downgrade or upgrade. The concrete test above is exactly the minimal check that would either justify elevating the novel indices or require them to be labeled prospective, which is what the conditional already asks for.","tokens_in":38386,"tokens_out":501,"duration_ms":10265,"concrete_test":"Pick one novel index (e.g., CFI from §7.8) and one classical metric (e.g., load-imbalance LID §4.13) on a public continuum trace or simulator run with known placement failures; check whether ranking nodes/configurations by CFI predicts placement success or SLO recovery better than LID alone. If it does not, mark the novel indices explicitly speculative in the camera-ready; if it does, the conditional is largely discharged.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is definitional and organizational: a three-layer taxonomy plus acquisition tags (scope/phase/method) and mathematical formulations for DCCS metrics, including several novel composites (CFI §7.8, Adaptivity Quotient §7.5, Observability E §7.4, Trustworthiness §7.15, etc.). For a taxonomy paper this is the appropriate epistemic standard; the formulations are mostly standard or transparent composites, internally consistent with the system model in §2, and the acquisition tables (2–5) are a genuine practical contribution. The reader's weakest assumption correctly flags that the novel indices lack empirical calibration and are not shown to change operator decisions. That is a real limitation on adoption value, but it does not undermine the soundness of the catalog itself, nor does it introduce an internal contradiction. No stronger load-bearing flaw (hidden inconsistency, misstated classical metrics, or circular reasoning) is present.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes a structured taxonomy of performance metrics for Distributed Computing Continuum Systems (DCCS), spanning cloud–fog–edge–mobile–IoT–sensor layers. Metrics are organized into computing-level, network-level, and application/user-level categories (Fig. 1; §§4–6), with mathematical formulations tied to a heterogeneous system model (§2). A distinctive contribution is the acquisition-requirement framing—scope (single-node / multi-node / full system / full app), phase (operational vs experimental), and method (standard telemetry / custom / experimental instrumentation)—summarized in Tables 2–5. Section 7 introduces emerging composite metrics aimed at sustainability, observability, adaptivity, data locality, continuum fragmentation, migration stability, and trustworthiness. The central claim is organizational and definitional: that this catalog, with formulations and acquisition tags, supports more transparent and consistent cross-layer DCCS evaluation than isolated single-dimension practices.","tokens_in":38657,"tokens_out":1737,"duration_ms":37445,"significance":"If adopted, the taxonomy would give the DCCS community a shared vocabulary and measurement checklist at a time when AI/edge workloads make single-layer benchmarks inadequate. The acquisition triple (scope/phase/method) is a practical contribution that goes beyond listing formulas and can guide monitoring design and experimental planning. Classical metrics are stated consistently with common usage and with the §2 model. The novel indices (CFI, Adaptivity Quotient, Observability score E, DLI, MSI, Trustworthiness, carbon/water efficiency) correctly name real continuum concerns that standard cloud or HPC suites under-emphasize. As a taxonomy paper, its value is catalog completeness and operational guidance rather than new theorems or empirical superiority proofs; on that standard the work is useful reference material for architects, benchmark designers, and continuum simulators (as the authors themselves flag in §8).","major_comments":[{"comment":"§7 (esp. §§7.4–7.5, 7.8–7.10, 7.15) and Table 5: Several novel composites (Observability E in Eq. (192), Adaptivity Quotient Q in Eq. (193), Continuum Fragmentation Index CFI in Eq. (196), Migration Stability Index MSI in Eq. (197), Adaptation Cost Efficiency ACE in Eq. (198), Trustworthiness TS_j in Eq. (208)) are introduced by definition only. There is no sensitivity analysis, toy example, or comparison against operator decisions or existing baselines showing that the indices discriminate continuum quality or change placement/orchestration choices. For a taxonomy this is acceptable if framed as proposals, but the Abstract and §1 currently present them as part of the evaluation solution on equal footing with classical metrics. Please either (i) add minimal worked examples / synthetic scenarios that show how each index moves under controlled continuum conditions, or (ii) clearly demote §","section":"§7; Table 5; Abstract; §1"},{"comment":"Free parameters without selection guidance undermine comparability—the stated goal of the taxonomy. Composite weights appear repeatedly (ya_i in Eq. (23), η_i in Eq. (44), a in congestion Eq. (106), yb_i in Eq. (174), w_i / α_r in flexibility and trustworthiness, M_max_k in MSI Eq. (197), AoI_max_i in freshness Eq. (207)). Section 3 discusses metric selection criteria (relevance, sensitivity, consistency, etc.) but does not address how to set or report these weights so that two studies remain comparable. Add a short subsection (likely under §3) on weight reporting conventions, recommended defaults, or normalization procedures, and note in §7 which novel metrics are parameter-free (e.g., CFI) versus weight-dependent.","section":"§3; Eqs. (23), (44), (106), (174), (197), (207), (208)"},{"comment":"Related-work positioning is thin relative to the claim of filling a gap in “transparent and consistent” DCCS evaluation. The introduction cites orchestration and offloading work but does not systematically contrast this taxonomy with existing cloud/edge/IoT benchmark suites, metric surveys, or standardization efforts (e.g., SPEC, TPC-style, EdgeBench-class tools, ISO/IEC sustainability metrics, or prior metric taxonomies in fog/edge). Without that map, it is hard for readers to see what is new beyond aggregation and the acquisition tags. A compact related-work or “gap analysis” subsection—even a table of prior taxonomies vs. coverage of continuum-specific dimensions (fragmentation, migration stability, data locality, AoI)—would make the contribution claim load-bearing rather than asserted.","section":"§1; before §2 or new related-work section"}],"minor_comments":[{"comment":"Duplicate/near-duplicate introductory paragraphs in §1 (the DCCS paradigm is described twice with overlapping cloud/edge/IoT wording). Tighten to a single narrative.","section":"§1"},{"comment":"Notation collisions and typos: Table 1 and §2 use both ι and mixed Greek/Latin; Eq. (7) averages RT_i while the left-hand side is TCT; “ya_1” style weights are unusual—prefer w_i or α_i consistently. “T able” line breaks in table captions; “CC F ragmentation”, “C. Efficiency”, “W ater”, “T rustworthiness” spacing artifacts in Table 5.","section":"Table 1; Eq. (7); Tables 2–5"},{"comment":"Fig. 1 is central but dense; consider splitting classical vs. novel branches or using a second figure for §7 metrics so acquisition tags can be visually linked.","section":"Fig. 1"},{"comment":"Amdahl/Gustafson appear in §4.3 and again conceptually in §7.6; cross-reference rather than re-motivate. Keywords list “Amdahl’s Law” though it is only one of many formulations—consider broader keywords (taxonomy, benchmarking, continuum metrics).","section":"§4.3; §7.6; Keywords"},{"comment":"Several acquisition rows mark Method as “Custom Instrumentation*” or Phase as “Experimental*” with asterisks explained only in prose; add a table footnote defining “*”.","section":"Tables 2–5"},{"comment":"§5.8 title is “Response Time” while the text says “Network response time” and Table 3 says “Network Response Time”—align names. §4.15 Cost Efficiency vs §6.2 Cost: briefly state how they differ to avoid double-counting in multi-layer studies.","section":"§5.8; §4.15; §6.2"},{"comment":"Future simulator claim in §8 is welcome; if any metric list, ontology, or machine-readable appendix is planned, mention format (e.g., whether formulations will be released as a catalog). Even a one-page checklist mapping evaluation goals → recommended metric subsets would increase immediate usability.","section":"§8; §3"}],"recommendation":"minor_revision","confidential_remarks":"Fit is appropriate for a metrics/taxonomy contribution in cs.DC. Overlap with the authors’ prior DCCS manifesto and orchestration papers is visible in motivation but does not appear to recycle metric identities. I would not block on lack of large-scale empirical validation if §7 is honestly framed as candidate metrics; I would block only if the authors insist the novel indices are ready-to-use evaluation standards without any demonstration. Minor revision is the right bar: three fixable structural points, not a rewrite of the catalog."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a careful taxonomy paper, not an empirical one. What you get is a three-layer organization (compute, network, app/user) under a clear continuum system model, standard formulas restated cleanly, and—actually useful—tables that tag each metric by acquisition scope, phase, and method. That last piece is the practical contribution: it tells an operator or experimenter whether something is single-node telemetry or full-system experimental instrumentation.\n\nMost of the classical metrics (latency, throughput, Amdahl/Gustafson, Jain fairness, AoI, EDP, SLO compliance) are standard and consistent with common usage. Section 3 on selection criteria and feasibility is sensible. The §7 composites—Continuum Fragmentation Index, Adaptivity Quotient, Observability score, Migration Stability Index, Carbon Efficiency per Useful Task, Trustworthiness—are named and formulaically transparent. They are not derived from first principles or checked against real systems, so treat them as proposals, not established instruments. Free weights and thresholds (η, α, M_max, etc.) are left open, which is honest but means the indices are not yet decision-ready.\n\nNo experiments, no baselines against prior metric surveys, no code. Circularity is low; self-citations mostly frame the DCCS setting rather than prop up the metric identities. Overlap with existing cloud/edge metric literature could be sharper, but that is a revision issue, not a soundness collapse.\n\nWho it is for: people writing DCCS evaluation sections, building simulators, or trying to stop apples-to-oranges reporting across edge–cloud papers. If you need a shared vocabulary and a feasibility checklist, this is worth having on the desk. If you need evidence that the new indices change operator decisions, it is not here yet.\n\nI would send it to referees. Ask them to push for clearer marking of unvalidated composites and tighter related-work positioning; the catalog itself holds up.","headline":"Solid DCCS metric catalog with useful acquisition tags; the novel composites are defined, not validated.","tokens_in":39326,"tokens_out":483,"would_cite":true,"duration_ms":16290,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A three-layer taxonomy with formulas and acquisition rules makes continuum performance comparable across cloud, edge, and IoT.","keywords":["distributed computing continuum","performance metrics","taxonomy","edge computing","metric acquisition","sustainability","observability","continuum fragmentation"],"falsifier":"Deploy the taxonomy’s novel indices on a multi-layer continuum testbed under controlled load, mobility, and failure scenarios and check whether they change scheduling or placement decisions and correlate with operator-visible outcomes better than classical metrics alone.","tokens_in":39240,"feed_emoji":"📐","tokens_out":843,"duration_ms":18194,"temperature":0.7,"pith_summary":"Distributed computing continuum systems spread work from sensors and phones up through edge and fog to the cloud, so no single metric can describe how well they run. This paper argues that evaluation has stayed fragmented—CPU here, latency there, energy somewhere else—and offers a structured taxonomy that groups metrics into computing-level, network-level, and application/user-level categories, then adds emerging ones for carbon, heat, observability, adaptivity, data locality, fragmentation, and migration stability. For representative metrics it supplies mathematical definitions and states where the data must come from (one node, many nodes, or the whole system), when (live operation or controlled experiment), and how (standard telemetry, custom code, or special instrumentation). The practical claim is that this shared map lets architects pick metrics that match their goals and compare architectures fairly instead of talking past one another.","feed_headline":"A taxonomy makes continuum performance comparable","feed_subtitle":"Formulas and acquisition rules span compute, network, apps, carbon, and fragmentation","key_machinery":"The taxonomy itself (Figure 1), plus per-metric acquisition triples (scope × phase × method) in Tables 2–5: these fix what is measured, how hard it is to collect, and whether it belongs in operations or experiments.","core_discovery":"Transparent, consistent evaluation of heterogeneous continuum systems requires a taxonomy that jointly organizes classical metrics (execution time, throughput, latency, packet loss, task success, SLO compliance, and so on) and novel continuum-specific ones (sustainability, observability, adaptivity, data locality, continuum fragmentation, migration stability), each with a formula and an acquisition profile of scope, phase, and method.","pith_inferences":["Without a reference open implementation of the novel indices, the taxonomy risks becoming a checklist rather than a working measurement standard.","The acquisition tables could double as a design filter: if a claimed ‘online adaptive’ controller needs Experimental Instrumentation and Full System scope, it is not yet ops-ready.","Carbon- and water-per-task metrics will pressure cloud and edge providers to expose footprint counters the way they already expose CPU and bandwidth.","Observability and trustworthiness scores point toward treating explainability and causal consistency as runtime QoS dimensions, not only offline audit traits."],"forward_implications":["Metric choice can be justified by acquisition cost: single-node standard telemetry for ops, multi-node or experimental instrumentation for design-time comparison.","Cross-paper continuum results become easier to compare when authors report the same layer-tagged metrics and formulas.","Sustainability (CO₂, heat, water per useful task) enters the same evaluation frame as latency and throughput.","Fragmentation, data locality, and migration stability become first-class checks for orchestration and placement algorithms.","Future continuum simulators can implement the listed formulas as a common measurement layer."],"fun_headline_variants":["Taxonomy unifies compute, network, and app metrics for continuum systems","Classical plus novel metrics make DCCS performance comparable","Formulas and acquisition profiles for continuum performance metrics","Organizing DCCS metrics across layers, sustainability, and fragmentation","Consistent continuum evaluation needs joint classical and novel metrics"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The new composite indices (fragmentation, adaptivity, observability, trustworthiness, and similar scores) are assumed to be well-defined and useful in real systems without empirical calibration against running deployments.","fun_headline_variants_meta":{"raw":{"variants":["Taxonomy unifies compute, network, and app metrics for continuum systems","Classical plus novel metrics make DCCS performance comparable","Formulas and acquisition profiles for continuum performance metrics","Organizing DCCS metrics across layers, sustainability, and fragmentation","Consistent continuum evaluation needs joint classical and novel metrics"]},"model":"grok-4.5","effort":"low","cost_usd":0.001848,"raw_usage":{"total_tokens":861,"prompt_tokens":775,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":18484000,"prompt_tokens_details":{"text_tokens":775,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":775,"tokens_out":66,"duration_ms":2649,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T08:19:24.248754+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Deploy the taxonomy’s novel indices on a multi-layer continuum testbed under controlled load, mobility, and failure scenarios and check whether they change scheduling or placement decisions and correlate with operator-visible outcomes better than classical metrics alone.","supporting_citations":[],"review_version":1}