REVIEW 2 major objections 1 minor 40 references
DECOFFEE lets independent edge nodes use reinforcement learning to decide workload offloading and cut delays, energy use, and dropped tasks.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-05-08 01:17 UTC
load-bearing objection DECOFFEE applies Double Dueling DQN with LSTM to decentralized offloading but the independent parallel MDP setup risks non-stationary learning that the simulations may not catch. the 2 major comments →
DECOFFEE: Decentralized Reinforcement Learning for Time-critical Workload Offloading and Energy Efficiency across the Computing Continuum
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
DECOFFEE formulates workload placement across the edge-cloud continuum as parallel Markov Decision Processes solved by autonomous agents running Double Dueling Deep Q-Networks augmented with LSTM-based load forecasting; each agent derives an offloading policy from local state and predicted conditions, and extensive simulations show these policies achieve lower system delay, lower energy consumption, and lower workload drop rates than conventional rule-based and heuristic strategies under varied traffic and network conditions.
What carries the argument
Multi-agent Double Dueling DQN with LSTM forecasting, where each independent edge agent learns an offloading policy from local observations to optimize joint delay-energy-drop objectives.
Load-bearing premise
The simulated task arrivals, node capacities, and changing link conditions are close enough to real edge-cloud settings that the policies learned by independent agents will perform similarly when deployed.
What would settle it
Deploy the trained DECOFFEE agents on a physical multi-node testbed with live IoT traffic generators and measure whether the reported reductions in delay, energy, and drops persist relative to the same heuristic baselines.
If this is right
- Edge nodes can reduce overall system delay by learning placement rules without a central coordinator sharing full state.
- Energy consumption across battery-constrained devices drops as agents avoid inefficient local or remote executions.
- Workload drop rate falls because agents anticipate load spikes and choose placements that respect timeout limits.
- The same learning structure maintains gains when traffic intensity or network variability increases.
Where Pith is reading between the lines
- The decentralized structure could extend to scenarios with node mobility or failures where central controllers become brittle.
- Replacing heuristic schedulers with local RL agents might simplify management in very large IoT deployments.
- Adding modest neighbor information sharing could further improve the policies without losing the no-coordination benefit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DECOFFEE, a decentralized multi-agent RL framework for time-critical workload offloading across the edge-cloud continuum. Workload placement is formulated as parallel Markov Decision Processes solved independently by Double Dueling DQN agents that use only local observations augmented by LSTM-based forecasts of future load. The central claim is that this approach jointly optimizes delay, energy consumption, and workload drop rate, with extensive simulations showing consistent outperformance over conventional rule-based and heuristic placement strategies under varying traffic and network conditions.
Significance. If the reported simulation gains are shown to be robust and not artifacts of the evaluation setup, the work would offer a scalable decentralized alternative to centralized optimization for latency-sensitive IoT workloads. The use of LSTM forecasting to handle time-varying conditions is a constructive design choice that addresses partial observability. No machine-checked proofs or parameter-free derivations are present, but the emphasis on fully decentralized execution without explicit inter-agent communication is a practical strength worth validating.
major comments (2)
- [Abstract and MDP formulation] Abstract and proposed framework: the workload placement is formulated as parallel independent Markov Decision Processes solved by independent Double Dueling DQN agents using only local observations. In reality, an agent's offload decision directly alters queue lengths, link congestion, and energy draw for neighboring nodes, rendering each agent's transition kernel dependent on the joint policy. This violates the Markov property for the individual MDPs and removes the convergence guarantees of standard DQN; the reported simulation improvements over baselines could therefore stem from a decoupled or sequential simulator rather than evidence that stable joint policies were learned. This issue is load-bearing for the paper's central claim of effective decentralized learning.
- [Simulation results] Simulation results section: the abstract states that DECOFFEE 'consistently outperform[s] conventional rule-based and heuristic placement strategies, achieving significant reductions' in delay, energy, and drop rate. However, the manuscript provides no description of the stochastic arrival process, heterogeneous node capacities, network trace sources, statistical significance testing, baseline re-implementations, or sensitivity analysis to DQN/LSTM hyperparameters. Without these details the support for the performance claims remains weak and the weakest assumption (that the simulated environment sufficiently represents real edge-cloud coupling) cannot be evaluated.
minor comments (1)
- [MDP formulation] Notation for the state, action, and reward definitions in the MDP formulation could be clarified with an explicit table or diagram to avoid ambiguity when readers compare the local observation space to the claimed global optimality.
Simulated Author's Rebuttal
We thank the referee for the constructive and detailed feedback on our manuscript. We have carefully considered each major comment and provide point-by-point responses below. Where appropriate, we have revised the manuscript to address the concerns and strengthen the presentation of our decentralized RL approach.
read point-by-point responses
-
Referee: [Abstract and MDP formulation] Abstract and proposed framework: the workload placement is formulated as parallel independent Markov Decision Processes solved by independent Double Dueling DQN agents using only local observations. In reality, an agent's offload decision directly alters queue lengths, link congestion, and energy draw for neighboring nodes, rendering each agent's transition kernel dependent on the joint policy. This violates the Markov property for the individual MDPs and removes the convergence guarantees of standard DQN; the reported simulation improvements over baselines could therefore stem from a decoupled or sequential simulator rather than evidence that stable joint policies were learned. This issue is load-bearing for the paper's central claim of effective decentralized learning.
Authors: We acknowledge the referee's valid observation that inter-agent dependencies exist in the real system, as offloading decisions affect shared resources. Our formulation intentionally uses independent MDPs with local observations only, augmented by LSTM-based load forecasting to approximate future states and mitigate partial observability. While this does not provide theoretical convergence guarantees for the joint policy (as would be the case in a fully centralized MDP), the approach is designed for practical decentralized execution without communication. The simulator implements coupled dynamics across nodes, and the empirical results demonstrate stable learning and consistent outperformance. In the revised manuscript, we have added a dedicated discussion subsection clarifying the approximation, its limitations relative to centralized methods, and why the LSTM component helps address non-Markovian effects in practice. revision: partial
-
Referee: [Simulation results] Simulation results section: the abstract states that DECOFFEE 'consistently outperform[s] conventional rule-based and heuristic placement strategies, achieving significant reductions' in delay, energy, and drop rate. However, the manuscript provides no description of the stochastic arrival process, heterogeneous node capacities, network trace sources, statistical significance testing, baseline re-implementations, or sensitivity analysis to DQN/LSTM hyperparameters. Without these details the support for the performance claims remains weak and the weakest assumption (that the simulated environment sufficiently represents real edge-cloud coupling) cannot be evaluated.
Authors: We appreciate the referee pointing out these omissions, which limit the reproducibility and strength of the evaluation claims. In the revised manuscript, we have substantially expanded the Simulation Setup and Results sections to include: (i) details of the stochastic arrival process (Poisson arrivals with time-varying rates drawn from real IoT workload traces), (ii) heterogeneous node capacities and energy models, (iii) sources of network traces (synthetic but calibrated to public edge-cloud datasets), (iv) statistical significance testing via paired t-tests (p < 0.05 reported for all key metrics), (v) explicit descriptions of how baselines were re-implemented, and (vi) sensitivity analysis across DQN and LSTM hyperparameters (learning rate, memory size, forecast horizon). These additions directly address the concern about the simulated environment's fidelity to real edge-cloud coupling. revision: yes
Circularity Check
No circularity in derivation chain
full rationale
The paper formulates workload placement as parallel MDPs solved via independent Double Dueling DQN agents augmented with LSTM forecasting, then reports empirical outperformance via simulations against rule-based and heuristic baselines. No load-bearing step reduces by construction to its own inputs: there are no self-definitional equations, no fitted parameters renamed as predictions, and no self-citation chains invoked to justify uniqueness or ansatzes. The simulation results constitute independent empirical evidence rather than a closed-form derivation that tautologically reproduces its assumptions.
Axiom & Free-Parameter Ledger
free parameters (1)
- DQN and LSTM hyperparameters
axioms (1)
- domain assumption Workload placement decisions can be modeled as parallel Markov Decision Processes solvable by independent agents
read the original abstract
The rapid proliferation of latency-sensitive and battery-constrained Internet-of-Things (IoT) applications has intensified the need for intelligent workload placement mechanisms across the Edge-Cloud computing continuum. In such environments, far-edge nodes must dynamically decide whether to execute workloads locally or offload them to neighboring nodes or the cloud, while accounting for execution delay, energy consumption, and strict timeout constraints. However, workload placement in large-scale distributed infrastructures is a highly dynamic and non-convex optimization problem due to stochastic arrivals, heterogeneous computing capacities, and time-varying network conditions. This paper proposes DECOFFEE, a decentralized reinforcement learning framework for time-critical workload offloading and energy-efficient operation across the computing continuum. The proposed multi-agent learning scheme jointly optimizes system delay, energy consumption, and workload drop rate through adaptive placement decisions. Each edge agent operates as an autonomous learning entity that derives an optimal policy from local system observations and predicted network conditions. The workload placement process is formulated as parallel Markov Decision Processes and solved using a Double Dueling Deep Q-Network (DQN) architecture enhanced with Long Short-Term Memory (LSTM) forecasting to anticipate future load conditions. Extensive simulations demonstrate that DECOFFEE and its variants consistently outperform conventional rule-based and heuristic placement strategies, achieving significant reductions in delay, energy consumption, and workload drop rate under varying traffic and network conditions.
Figures
Reference graph
Works this paper leans on
-
[1]
The computing continuum: From iot to the cloud,
A. Al-Dulaimy, M. Jansen, B. Johansson, A. Trivedi, A. Iosup, M. Ashjaei, A. Galletta, D. Kimovski, R. Prodan, K. Tserpes et al., “The computing continuum: From iot to the cloud,” Internet of Things, vol. 27, p. 101272, 2024
work page 2024
-
[2]
Placing Computational Tasks Within Edge-Cloud Continuum: A DRL Delay Minimization Scheme,
A. Giannopoulos, A. L. Suárez-Cetrulo, X. Masip-Bruin, F. D’Andria, and P. Trakadas, “Placing Computational Tasks Within Edge-Cloud Continuum: A DRL Delay Minimization Scheme,” in European Conference on Parallel Processing. Springer, 2024, pp. 36–45
work page 2024
-
[3]
Emerging edge computing technologies for distributed iot systems,
A. Alnoman, S. K. Sharma, W. Ejaz, and A. Anpalagan, “Emerging edge computing technologies for distributed iot systems,” IEEE Network, vol. 33, no. 6, pp. 140–147, 2019
work page 2019
-
[4]
Learning anticipatory decision for distributed systems with robustness guarantees,
P. Liu, X. Yang, H. Ren, H. Zhang, and Z. Wang, “Learning anticipatory decision for distributed systems with robustness guarantees,” IEEE Transactions on Automation Science and Engineering, 2024
work page 2024
-
[5]
Exploring the potential of distributed computing continuum systems,
P. K. Donta, I. Murturi, V. Casamayor Pujol, B. Sedlak, and S. Dustdar, “Exploring the potential of distributed computing continuum systems,” Computers, vol. 12, no. 10, p. 198, 2023
work page 2023
-
[6]
A. Giannopoulos, I. Paralikas, S. Spantideas, and P. Trakadas, “HOODIE: Hybrid computation offloading via distributed deep reinforcement learning in delay-aware cloud-edge continuum,” IEEE Open Journal of the Communications Society, 2024
work page 2024
-
[7]
A review on computational intelligence techniques in cloud and edge computing,
M. Asim, Y. Wang, K. Wang, and P.-Q. Huang, “A review on computational intelligence techniques in cloud and edge computing,” IEEE Transactions on Emerging Topics in Com- putational Intelligence, vol. 4, no. 6, pp. 742–763, 2020
work page 2020
-
[8]
Dynamic service placement for mobile micro-clouds with predicted future costs,
S. Wang, R. Urgaonkar, T. He, K. Chan, M. Zafer, and K. K. Leung, “Dynamic service placement for mobile micro-clouds with predicted future costs,” IEEE Transactions on Parallel and Distributed Systems, vol. 28, no. 4, pp. 1002–1016, 2016
work page 2016
-
[9]
Deadline-constrained multi-resource task mapping and allocation for edge-cloud sys- tems,
C. Gao, A. Shaan, and A. Easwaran, “Deadline-constrained multi-resource task mapping and allocation for edge-cloud sys- tems,” in GLOBECOM 2022-2022 IEEE Global Communica- tions Conference. IEEE, 2022, pp. 5037–5043
work page 2022
-
[10]
D. K. Nishad, V. R. Verma, P. Rajput, S. Gupta, A. Dwivedi, and D. R. Shah, “Adaptive ai-enhanced computation offloading with machine learning for qoe optimization and energy-efficient mobile edge systems,” Scientific Reports, vol. 15, no. 1, p. 15263, 2025
work page 2025
-
[11]
Computation offloading in resource-constrained multi-access edge computing,
K. Li, X. Wang, Q. He, J. Wang, J. Li, S. Zhan, G. Lu, and S. Dustdar, “Computation offloading in resource-constrained multi-access edge computing,” IEEE Transactions on Mobile Computing, vol. 23, no. 11, pp. 10 665–10 677, 2024
work page 2024
-
[12]
A. Giannopoulos, I. Paralikas, S. Spantideas, N. Nomikos, and P. Trakadas, “PDPPnet: Prioritized Delay-aware and Peer-to- Peer Task Offloading in Cloud-Edge Continuum with Double Dueling Deep Q-Networks,” in 2024 IEEE 29th International Workshop on Computer Aided Modeling and Design of Com- munication Links and Networks (CAMAD). IEEE, 2024, pp. 1–8
work page 2024
-
[13]
Resource scheduling in edge computing: A survey,
Q. Luo, S. Hu, C. Li, G. Li, and W. Shi, “Resource scheduling in edge computing: A survey,” IEEE communications surveys & tutorials, vol. 23, no. 4, pp. 2131–2165, 2021
work page 2021
-
[14]
E. Moro and I. Filippini, “Joint management of compute and radio resources in mobile edge computing: A market equilibrium approach,” IEEE Transactions on Mobile Computing, vol. 22, no. 2, pp. 983–995, 2021
work page 2021
-
[15]
An efficient distributed task offloading scheme for vehicular edge computing networks,
M. S. Bute, P. Fan, L. Zhang, and F. Abbas, “An efficient distributed task offloading scheme for vehicular edge computing networks,” IEEE Transactions on Vehicular Technology, vol. 70, no. 12, pp. 13 149–13 161, 2021
work page 2021
-
[16]
Dependent task offloading for edge computing based on deep reinforcement learning,
J. Wang, J. Hu, G. Min, W. Zhan, A. Y. Zomaya, and N. Georgalas, “Dependent task offloading for edge computing based on deep reinforcement learning,” IEEE Transactions on Computers, vol. 71, no. 10, pp. 2449–2461, 2021
work page 2021
-
[17]
Prioritization based task offloading in UA V-assisted edge networks,
O. Kalinagac, G. Gür, and F. Alagöz, “Prioritization based task offloading in UA V-assisted edge networks,” Sensors, vol. 23, no. 5, p. 2375, 2023
work page 2023
-
[18]
Deep reinforcement learning for task offloading in mobile edge computing systems,
M. Tang and V. W. Wong, “Deep reinforcement learning for task offloading in mobile edge computing systems,” IEEE Transactions on Mobile Computing, vol. 21, no. 6, pp. 1985– 1997, 2020
work page 1985
-
[19]
Deep reinforcement learning techniques for dynamic task of- floading in the 5g edge-cloud continuum,
G. Nieto, I. De la Iglesia, U. Lopez-Novoa, and C. Perfecto, “Deep reinforcement learning techniques for dynamic task of- floading in the 5g edge-cloud continuum,” Journal of Cloud Computing, vol. 13, no. 1, p. 94, 2024
work page 2024
-
[20]
Edge intelli- gence for energy-efficient computation offloading and resource allocation in 5g beyond,
Y. Dai, K. Zhang, S. Maharjan, and Y. Zhang, “Edge intelli- gence for energy-efficient computation offloading and resource allocation in 5g beyond,” IEEE Transactions on Vehicular Technology, vol. 69, no. 10, pp. 12 175–12 186, 2020
work page 2020
-
[21]
N. Yang, S. Chen, H. Zhang, and R. Berry, “Beyond the edge: An advanced exploration of reinforcement learning for mobile edge computing, its applications, and future research trajectories,” IEEE Communications Surveys & Tutorials, vol. 27, no. 1, pp. 546–594, 2024
work page 2024
-
[22]
F. G. Wakgra, B. Kar, S. B. Tadele, S.-H. Shen, and A. U. Khan, “Multi-objective offloading optimization in mec and vehicular- fog systems: A distributed-td3 approach,” IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 11, pp. 16 897– 16 909, 2024
work page 2024
-
[23]
Deepedge: A deep reinforcement learning based task orchestrator for edge computing,
B. Yamansavascilar, A. C. Baktir, C. Sonmez, A. Ozgovde, and C. Ersoy, “Deepedge: A deep reinforcement learning based task orchestrator for edge computing,” IEEE Transactions on Network Science and Engineering, vol. 10, no. 1, pp. 538–552, 2022
work page 2022
-
[24]
A. Giannopoulos, I. Paralikas, S. Spantideas, and P. Trakadas, “COOLER: Cooperative Computation Offloading in Edge- Cloud Continuum Under Latency Constraints via Multi-Agent Deep Reinforcement Learning,” in 2024 International Confer- ence on Intelligent Computing, Communication, Networking and Services (ICCNS). IEEE, 2024, pp. 9–16
work page 2024
-
[25]
Bottleneck identification in cloudified mobile networks based on distributed telemetry,
M.-R. Fida, A. H. Ahmed, T. Dreibholz, A. F. Ocampo, A. Elmokashfi, and F. I. Michelinakis, “Bottleneck identification in cloudified mobile networks based on distributed telemetry,” IEEE Transactions on Mobile Computing, vol. 23, no. 5, pp. 5660–5676, 2023
work page 2023
-
[26]
Towards a lightweight distributed telemetry for microservices,
M. Otero, J. M. Garcia, and P. Fernandez, “Towards a lightweight distributed telemetry for microservices,” in 2024 IEEE 44th International Conference on Distributed Computing Systems Workshops (ICDCSW). IEEE, 2024, pp. 75–82
work page 2024
-
[27]
Delay-optimal computation task scheduling for mobile-edge computing sys- tems,
J. Liu, Y. Mao, J. Zhang, and K. B. Letaief, “Delay-optimal computation task scheduling for mobile-edge computing sys- tems,” in 2016 IEEE international symposium on information theory (ISIT). IEEE, 2016, pp. 1451–1455
work page 2016
-
[28]
Statistical analysis of the generalized processor sharing scheduling discipline,
Z.-L. Zhang, D. Towsley, and J. Kurose, “Statistical analysis of the generalized processor sharing scheduling discipline,” IEEE Journal on Selected Areas in Communications, vol. 13, no. 6, pp. 1071–1080, 1995
work page 1995
-
[29]
Y. Chen, N. Zhang, Y. Zhang, X. Chen, W. Wu, and X. S. Shen, “Toffee: Task offloading and frequency scaling for energy efficiency of mobile devices in mobile edge computing,” IEEE Transactions on Cloud Computing, vol. 9, no. 4, pp. 1634–1644, 2019
work page 2019
-
[30]
Online management for edge-cloud collaborative continuous learning: A two-timescale approach,
S. Lin, X. Zhang, Y. Li, C. Joe-Wong, J. Duan, D. Yu, Y. Wu, and X. Chen, “Online management for edge-cloud collaborative continuous learning: A two-timescale approach,” IEEE Transactions on Mobile Computing, vol. 23, no. 12, pp. 14 561–14 574, 2024
work page 2024
-
[31]
Multi-agent drl-based task offloading in multiple ris-aided iov networks,
B. Hazarika, K. Singh, S. Biswas, S. Mumtaz, and C.-P. Li, “Multi-agent drl-based task offloading in multiple ris-aided iov networks,” IEEE Transactions on Vehicular Technology, vol. 73, no. 1, pp. 1175–1190, 2023
work page 2023
-
[32]
Deep reinforcement learning: From q-learning to deep q-learning,
F. Tan, P. Yan, and X. Guan, “Deep reinforcement learning: From q-learning to deep q-learning,” in International Conference on Neural Information Processing. Springer, 2017, pp. 475–483
work page 2017
-
[33]
The bellman equation for minimizing the maximum cost
E. Barron and H. Ishii, “The bellman equation for minimizing the maximum cost. ” Nonlinear Anal. Theory Methods Applic., vol. 13, no. 9, pp. 1067–1090, 1989
work page 1989
-
[34]
Experience replay for continual learning,
D. Rolnick, A. Ahuja, J. Schwarz, T. Lillicrap, and G. Wayne, “Experience replay for continual learning,” Advances in neural information processing systems, vol. 32, 2019. PREPRINT 22
work page 2019
-
[35]
M. Shokrnezhad, T. Taleb, and P. Dazzi, “Double deep q- learning-based path selection and service placement for latency- sensitive beyond 5g applications,” IEEE Transactions on Mobile Computing, vol. 23, no. 5, pp. 5097–5110, 2023
work page 2023
-
[36]
C. S. Pabla, “Completely fair scheduler,” Linux Journal, vol. 2009, no. 184, p. 4, 2009
work page 2009
-
[37]
C. Wang, C. Liang, F. R. Yu, Q. Chen, and L. Tang, “Com- putation offloading and resource allocation in wireless cellular networks with mobile edge computing,” IEEE Transactions on Wireless Communications, vol. 16, no. 8, pp. 4924–4938, 2017
work page 2017
-
[38]
Enhanced round-robin algorithm in the cloud computing environment for optimal task scheduling,
F. Alhaidari and T. Z. Balharith, “Enhanced round-robin algorithm in the cloud computing environment for optimal task scheduling,” Computers, vol. 10, no. 5, p. 63, 2021
work page 2021
-
[39]
Offloading schemes in mobile edge computing for ultra-reliable low latency communications,
J. Liu and Q. Zhang, “Offloading schemes in mobile edge computing for ultra-reliable low latency communications,” Ieee Access, vol. 6, pp. 12 825–12 837, 2018. ANASTASIOS E. GIANNOPOULOS (Mem- ber, IEEE) (M.Eng, Ph.D) received the diploma of Electrical and Computer Engineer- ing from the National Technical University of Athens (NTUA), where he also comple...
work page 2018
-
[40]
Development of Methods for obtaining DC and low frequency AC magnetic cleanliness in space missions
He also obtained his Ph.D. at the Wireless and Long Distance Communications Laboratory of NTUA. His research interests include advanced Optimization Techniques for Wireless Systems, ML-assisted Resource Al- location, Maritime Communications and Multi-dimensional Data Analysis. He is currently working as a Research Associate at the National and Kapodistria...
work page 2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.