{"id":"2fdad901-edf5-4905-83bd-ed1c3cf31ed5","arxiv_id":"2502.06963","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature survey classifies DRL-based vehicular task offloading into centralized, distributed, and hierarchical architectures and identifies recurring gaps in MDP modeling, reward design, and evaluation.","lead":"This paper is a survey of deep reinforcement learning (DRL) methods for deciding where vehicle tasks should be computed in edge computing networks. It organizes the literature into centralized, distributed, and hierarchical offloading architectures and lists common weaknesses such as incomplete state models and weak reward design.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's aggregate lessons rest on unverifiable classifications and undisclosed paper selection; internal mischaracterizations (e.g., ref [75] in §5.3 vs Table 5) show the taxonomy cannot be checked from the manuscript.","rationale":"The paper is a survey, so its evidence base is its citations. The single most load-bearing assumption is that the selection and coding of those citations are accurate and representative. The manuscript gives no methodology for selection, and the internal classification errors I found are concrete evidence that coding is not reliable. Therefore, the aggregate lessons cannot be accepted as systematic findings without an audit. This does not negate the survey's utility as an organized map, but it means conditional acceptance with mandatory disclosure of methodology and a corrected, auditable classification table is the right bar. I agree with the reader's weakest-assumption identification. I credit the paper for a broad citation base and a clear organizational structure, but those do not substitute for verifiable coding. The proposed re-coding test would settle whether the concern lands by measuring agreement between independent labels and the paper's tables and by testing representativeness of the 133-paper set.","tokens_in":31856,"tokens_out":4119,"duration_ms":36997,"concrete_test":"Independently re-code a stratified random sample of 20 papers from Tables 5-7 (at least 6 per table): for each, record architecture (centralized/distributed/hierarchical), DRL algorithm, whether state includes channel/task/resource dynamics, whether reward includes transmission energy and multi-objective weights, and which baselines are compared. Compare against the paper's Table entries and §5.5/6.5/7.5 claims. If agreement is below 90%, or if the sampled papers' MDP/reward/baseline patterns differ materially from the aggregate lessons, the central conclusions do not generalize. As a second check, run a reproducible Scopus/Web-of-Science search with explicit VEC+DRL+offloading terms and year range, and test whether the 133-paper set is a representative subset; report inclusion/exclusion decisions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Central claim: the three-way taxonomy and lessons about incomplete MDPs, miscalibrated rewards, and weak baselines. These depend on (a) the 133 papers being representative and (b) the authors' classifications being accurate. Neither is supported. Section 1.3 claims 'comprehensive' with no search strategy, inclusion criteria, databases, or date range. Internal coding is demonstrably unreliable: [75] is Table 5 'MATD3 (DTDE)' yet §5.3 describes it as PPO and as MADDPG/MAPPO CTDE; [76] is centralized in Table 5 but 'fully decentralized' in §6.3. Since Section 8's lessons are aggregations over these coded papers, misclassification propagates into every general conclusion. The promised MDP analysis is also not delivered: no systematic coding of state/action/reward across the corpus, so it is unverifiable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews deep reinforcement learning (DRL) approaches to computational offloading in vehicular edge computing (VEC). It organizes roughly 133 cited papers into three architectural categories—centralized, distributed, and hierarchical—and, for each, summarizes DRL methods, optimization objectives, key techniques, and reported limitations. The paper also provides background on MDPs, DRL algorithms, and multi-agent RL, and ends with lessons learned and open research directions centered on incomplete MDP representations, miscalibrated reward functions, weak baselines, synchronization issues, and scalability.","tokens_in":31932,"tokens_out":6442,"duration_ms":57757,"significance":"If the survey's classifications and aggregate lessons are accurate, it provides a useful structured map of an active and crowded area, and its emphasis on MDP completeness, reward scaling, and evaluation practice is a plausible synthesis of recurring problems. The paper's strengths are its breadth of coverage, its per-architecture limitation subsections, and its attention to action-space and coordination issues that are often treated superficially in primary papers. However, the contribution is currently weakened by the absence of a disclosed literature-search methodology, by internal inconsistencies in the coding of individual papers, and by the fact that the promised systematic MDP analysis is not actually presented. The paper does not ship code or machine-checked artifacts, so its value rests entirely on the reliability of its synthesis.","major_comments":[{"comment":"The paper claims to be a 'comprehensive review' and to 'systematically investigate' DRL-based offloading, but it never specifies the search strategy: no databases, keywords, inclusion/exclusion criteria, date range, or screening process are given. Since Section 8's lessons are aggregate statements over the selected 133 papers, the representativeness of that corpus is load-bearing. Without this methodology, the reader cannot distinguish a systematic survey from a convenience sample.","section":"Abstract and §1.3"},{"comment":"The coding of individual papers is internally inconsistent, which directly threatens the taxonomy on which the survey's conclusions rest. In Table 5, reference [75] is classified as 'MATD3 (DTDE)', yet §5.3 states that PPO is applied in [75] and that MADDPG and MAPPO are used in [73, 75]. Reference [76] appears in Table 5 as a centralized MAAC (DTDE) scheme, but §6.3 lists [76] among 'fully decentralized variants of MADQN and MAAC'. Additionally, the same critique of [68] appears in both §5.5.1 and §6.5.1, despite [68] being coded as a centralized scheme. The authors should state explicit, mutually exclusive classification criteria for 'centralized', 'distributed', and 'hierarchical', and re-code all entries consistently; otherwise the aggregate lessons in Section 8 cannot be trusted.","section":"Tables 5–7, §5.3, §6.3"},{"comment":"The paper promises an analysis of MDP formulations—'we analyze how Markov Decision Process (MDP) formulations are applied'—and a contribution bullet claims to 'assess the variability and limitations of reward function design' and 'examine the role of action space representation'. However, no systematic coding of state, action, transition, and reward components across the surveyed papers is provided. Tables 5–7 list DRL method, optimization objective, key technique, and computing source, but not the MDP tuple elements. Consequently, statements such as those in §8.1 about 'incomplete state representations' are qualitative assertions rather than results of a documented cross-corpus analysis. A dedicated MDP-coding table or appendix is needed to substantiate the paper's central claims.","section":"Abstract and §1.3 vs. Tables 5–7 and §8.1"},{"comment":"The per-paper critiques that motivate the survey's open-problem lessons are not verifiable from the manuscript. For example, §5.5.1 says [69] 'fails to incorporate task and resource dynamics into the state representation', and §7.5.1 says [108] 'neglected to include task dynamics, resource dynamics, and channel variability', but the authors do not quote the original state vectors, reward equations, or experimental configurations. Since these critiques are the basis for the survey's conclusions about incomplete MDPs and miscalibrated rewards, the authors should cite specific sections or equations from the original papers, or provide a per-paper MDP coding matrix that allows the reader to check each judgment.","section":"§5.5, §6.5, §7.5"}],"minor_comments":[{"comment":"The sentence beginning 'Focusing on One of the key challenges that arise in vehicular networks...' is grammatically incomplete, and nearby phrasing such as 'Surveys such as those [16] and [17] emphasized' and 'the work of [7]' is awkward. Please revise this paragraph for clarity.","section":"§1.2"},{"comment":"The captions of Figures 7 and 8 contain the typo 'Hierarchial'; this should be 'Hierarchical'.","section":"Figures 7 and 8"},{"comment":"The text calls MAPPO a deterministic policy, but PPO and its multi-agent variant are stochastic on-policy methods, as correctly stated in Section 4.3.3 and Table 3. Please correct this inconsistency.","section":"§6.5.3"},{"comment":"The MAAC algorithm is cited to [62], which is Zhang et al.'s 'Fully Decentralized Multi-Agent Reinforcement Learning With Networked Agents'; this reference does not describe the MAAC (multi-agent actor-critic with a centralized critic conditioned on other agents' actions) as the term is used in the MARL literature. Please update the citation.","section":"§4.4.3"},{"comment":"DQN, Double DQN, and Dueling DQN are labeled 'Deterministic' in the policy column. While the learned greedy policy is deterministic, the behavior policy during exploration is stochastic (e.g., epsilon-greedy). Consider clarifying this distinction in the table or its notes.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The self-citation [81] is one of many references and does not by itself distort the survey's conclusions. The main risk is that the aggregate lessons are only as trustworthy as the unshown coding of the 133 papers; if the authors cannot supply the MDP coding matrix and reconcile the inconsistencies in Tables 5–7, the survey's central claims about incomplete MDPs, reward design, and evaluation gaps would remain unsubstantiated. I therefore support a major revision rather than rejection, because the taxonomy and lessons are defensible in principle and can be repaired within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a survey, not a research contribution. It organizes roughly 130 recent papers on DRL-based task offloading in vehicular edge computing into centralized, distributed, and hierarchical architectures, with per-paper critiques and comparison tables. That is genuinely useful: the subfield is crowded, and this gives a newcomer a quick map of who did what, with which algorithm, for which objective. The specific critiques in Sections 5.5, 6.5, and 7.5 (incomplete MDP state space, miscalibrated rewards, weak baselines) are concrete and will save people time. Section 8's lessons are sensible, if not original.\n\nThe problems are real. The classifications are internally inconsistent in ways you can verify without leaving the manuscript. Reference [75] is listed as MATD3 (DTDE) in Table 5, but Section 5.3 describes it as PPO in one place and MADDPG/MAPPO under CTDE in another. Reference [76] is centralized in Table 5 but called 'fully decentralized' in Section 6.3. If the authors cannot align their own text and tables, the accuracy of their characterizations of other papers is doubtful. The abstract and Section 1.3 promise a systematic analysis of MDP formulations, but no such analysis appears: no per-paper breakdown of state, action, reward, or transition components. The claim of 'comprehensive' coverage is also unsupported by any disclosed search strategy, inclusion criteria, or date range. These are not minor stylistic issues; they undermine the survey's credibility as a reliable index.\n\nOn the positive side, the citation pattern looks honest. There is one self-citation ([81]), but the taxonomy does not depend on it, and it appears to be a legitimate related work. The reference list spans 2019–2024 and seems representative, though without a methodology we cannot verify that.\n\nMy take: the high-level narrative—that DRL offloading solutions cluster into three architectures and share recurring weaknesses—holds up as a framing device. But the specific coding errors mean the aggregate lessons, which are built on those classifications, are not fully trustworthy. A reader should treat the paper as a starting point and check any specific claim about a particular cited work against the original source.\n\nSend it to peer review. The paper deserves referee time because it is a useful map and the high-level lessons are sound. The referee should require the authors to fix the internal inconsistencies, disclose their review methodology, and either deliver the MDP analysis or drop the claim. With those revisions it could be a solid survey; as it stands, it is a draft that a careful editor would send back for revision rather than accept as-is.","headline":"A useful survey map of DRL-based offloading in vehicular edge computing, but the internal classification inconsistencies and the unfulfilled MDP analysis claim make it a resource to use with caution.","tokens_in":32510,"tokens_out":2917,"would_cite":true,"duration_ms":27617,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that DRL-based offloading in vehicular edge computing sorts into centralized, distributed, and hierarchical architectures, and that the field's recurring bottlenecks are incomplete MDP state representations…","keywords":["vehicular edge computing","deep reinforcement learning","computation offloading","Markov decision process","multi-agent reinforcement learning","reward function design","centralized distributed hierarchical architectures","intelligent transportation systems"],"falsifier":"A reader could settle the core claim by taking the papers the survey places in each architecture and re-coding each one's state variables, reward terms, and baseline comparisons against a fixed checklist; if the alleged omissions, such as missing transmission energy, missing channel dynamics, or absent DRL baselines, appear in a different pattern or mostly disappear, the survey's recurring-limitation thesis would fail.","tokens_in":31594,"feed_emoji":"🚗","tokens_out":5058,"duration_ms":45837,"temperature":0.7,"pith_summary":"The paper is a survey that tries to establish that deep reinforcement learning (DRL) has become the main adaptive tool for deciding where vehicles send computational tasks, and that the resulting literature is best organized into centralized, distributed, and hierarchical offloading architectures. It claims that across all three architectures the same recurring defects appear: state representations that omit important dynamics, reward functions that are miscalibrated for multi-objective goals, and evaluations that lack proper DRL baselines. A sympathetic reader would take the intended contribution as a map of the field's design space plus a diagnosis of what blocks real deployment. If the diagnosis holds, future progress depends less on inventing new DRL algorithms and more on making MDP formulations complete, rewards well scaled, and comparisons honest.","feed_headline":"Survey: DRL vehicular offloading splits into three architectures","feed_subtitle":"Centralized, distributed, and hierarchical designs share MDP-state, reward-scaling, and baseline gaps.","key_machinery":"The central organizing device is the Markov Decision Process (MDP), the tuple $(S,A,P,R,\\gamma)$ that defines state, action, transition, reward, and discount. The survey uses MDP/POMDP formulation as a lens to classify each offloading scheme and to expose omissions: which dynamics are left out of the state, how the reward weights competing objectives, and whether multi-agent learning assumes centralized training or decentralized execution. The architecture taxonomy (centralized, distributed, hierarchical) is the second load-bearing structure; it tells readers what kind of coordination problem each paper solves.","core_discovery":"On its own terms, the survey's central claim is that DRL-based offloading in vehicular edge computing matures through three architectural paradigms, each with its own MDP structure: centralized offloading with a single agent or multiple agents reporting to one server; distributed offloading where multiple edge servers learn jointly, often under centralized training with decentralized execution; and hierarchical offloading across vehicles, edge, fog, UAVs, and cloud. It further claims that the field's open problems are not primarily algorithmic novelty but MDP completeness, reward-function calibration, baseline comparison, synchronization, and scalability. The evidence for this claim is a structured reading of roughly 130 studies, organized into tables that pair each method with its DRL algorithm, optimization objective, and computing source, followed by a section-by-section list of limitations that recur across independent works.","pith_inferences":["If the recurring-limitation pattern is real, many reported gains over weak baselines may shrink when compared against a strong, fully informed MDP baseline, so public benchmark suites with fixed state and reward definitions would be the natural next step.","The survey's emphasis on incomplete state representations suggests that offline RL or model-based RL, which can leverage logged data from real vehicular systems, might progress faster than simulation-only DRL.","Because most surveyed studies rely on small synthetic simulations, hardware-in-the-loop or field trials could overturn specific performance claims even if the taxonomy survives.","Co-adaptive game-theory-plus-DRL designs, where policies update payoff models and equilibria co-evolve, are a plausible direction the survey names as open rather than a proven remedy."],"forward_implications":["If the survey's diagnosis is right, proposed offloading policies should be evaluated on whether their state includes transmission energy, channel dynamics, mobility, queue length, and caching or task dependencies.","Reward design should move from raw or squared latency and energy terms to scaled, weighted, windowed objectives so that no single metric dominates learning.","Performance claims need consistent DRL baselines, including multi-agent baselines, plus QoS and QoE metrics rather than utility-only curves.","Multi-agent systems should adopt synchronization and coordination mechanisms, and new agents should be able to join without full retraining.","Deployment-focused work should address communication delays, computational constraints, and high-dimensional action spaces, with lightweight or generalized models."],"supporting_citations":[{"why":"Supplies the MEC-from-the-communication-perspective baseline that motivates edge offloading in vehicular networks.","marker":"[3]"},{"why":"Establishes the prior survey of RL/DRL for vehicular task offloading that this survey positions itself against and extends.","marker":"[4]"},{"why":"Provides the earlier MEC survey whose scope, covering latency reduction, resource allocation, and privacy, frames the edge-computing background.","marker":"[16]"},{"why":"Defines edge computing's vision and challenges, grounding the hierarchical multi-tier discussion.","marker":"[17]"},{"why":"Documents MDP limitations in IoT DRL and motivates the survey's call for robust decentralized and multi-agent frameworks.","marker":"[25]"},{"why":"Surveys multi-agent RL in future Internet technologies and supplies the MARL coordination background used in the later sections.","marker":"[28]"},{"why":"Provides the MDP and policy-gradient definitions on which the survey's analytical frame rests.","marker":"[40]"}],"fun_headline_variants":["DRL offloading review: three architectures, shared gaps","Vehicular DRL offloading: centralized, distributed, hierarchical","DRL vehicular offloading survey: three paradigms, common flaws","Review: DRL offloading in VEC has three architecture types, shared pitfalls"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the roughly 130 papers the survey selected and sorted into centralized, distributed, and hierarchical architectures fairly represent the field, so its lessons about MDP completeness, reward design, and evaluation gaps generalize.","fun_headline_variants_meta":{"raw":{"variants":["DRL offloading review: three architectures, shared gaps","Vehicular DRL offloading: centralized, distributed, hierarchical","DRL vehicular offloading survey: three paradigms, common flaws","Review: DRL offloading in VEC has three architecture types, shared pitfalls"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000572,"raw_usage":{"total_tokens":2673,"prompt_tokens":883,"completion_tokens":1790,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":1714}},"tokens_in":499,"tokens_out":1790,"duration_ms":11551,"temperature":1.0,"reasoning_tokens":1714,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T14:12:07.435219+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could settle the core claim by taking the papers the survey places in each architecture and re-coding each one's state variables, reward terms, and baseline comparisons against a fixed checklist; if the alleged omissions, such as missing transmission energy, missing channel dynamics, or absent DRL baselines, appear in a different pattern or mostly disappear, the survey's recurring-limitation thesis would fail.","supporting_citations":[],"review_version":1}