{"id":"96b87099-25f5-49c0-885e-5304aa09fd96","arxiv_id":"2502.09417","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured survey of reinforcement learning methods applied to optimization across manufacturing, energy, and robotics, with challenges and future directions.","lead":"This paper reviews how reinforcement learning is used to solve optimization problems in manufacturing, energy systems, and robotics. It organizes dozens of recent studies into tables and discusses common challenges like sample efficiency, safety, and real-world deployment.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table II lists four HVAC-control papers (refs 29, 47, 48, 45) as the RL methods for Renewable Energy Integration, so the survey's core categorization is internally inconsistent and needs verification.","rationale":"I read the paper as a survey whose contribution is organizational. Its strongest claim is firstness and comprehensiveness. The reader's weakest assumption is the lack of a systematic search methodology; my concern is stronger and more concrete: the internal evidence already shows the tables are unreliable. A reader cannot verify representativeness without knowing the search protocol, but anyone can verify that the Renewable Energy Integration row lists HVAC papers. This makes the central 'systematic categorization' claim questionable as written. The correct remedy is a careful audit and correction of the tables, so conditional acceptance remains appropriate. I do not see grounds to reject because the narrative sections are coherent and the paper can be repaired.","tokens_in":13816,"tokens_out":6529,"duration_ms":52152,"concrete_test":"One check: read the abstracts of every reference in Table II and assign each to the subdomain it actually addresses, then compare with the table. The specific decisive case is whether PPO [29], Munchausen DQN [47], Q-learning [48], and A3C [45] belong under Renewable Energy Integration; they are HVAC papers, so the row is likely shifted. If confirmed, audit Tables I, III, and IV for the same misalignment pattern before the survey is described as comprehensive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it provides the first survey specifically focused on RL for optimization in automation, with a systematic categorization and a comprehensive bibliography (Section I.B). For this claim to hold, the application tables must at least assign cited works to the correct subdomains. Table II does not: the 'Renewable Energy Integration' column lists PPO [29], Batch Constrained Munchausen Deep Q-learning [47], Q-learning [48], and A3C [45] as its RL approaches. Checking the references shows these are HVAC-control papers: [29] is whole-building HVAC and demand response, [47] is safe HVAC control via batch RL, [48] is HVAC operation optimization, and [45] is end-to-end DRL for HVAC in office buildings. None is a renewable-energy-integration method. The narrative text for renewable integration instead cites [4], [31], [41]-[44], so the table contradicts the paper's own prose. This is a verifiable internal inconsistency, not a criticism that depends on an external search protocol. Because Section II's tables are the main deliverable of the survey, this error calls the reliability of the 'systematic categorization' into question until a full audit is performed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of reinforcement learning (RL) applied to optimization problems in automation, organized around three application domains: manufacturing, energy systems, and robotics. It provides comparative tables (Tables I–III) that categorize representative studies by sub-domain, RL approach, methodology, outcomes, and future directions. Section III discusses five cross-cutting challenges—sample efficiency and scalability; safety and robustness; interpretability and trustworthiness; transfer learning and meta-learning; and real-world deployment and integration—and Section IV offers concluding remarks. The authors claim this is the first survey specifically focused on RL for optimization in automation and that it provides a systematic categorization and comprehensive bibliography.","tokens_in":14023,"tokens_out":3810,"duration_ms":32207,"significance":"If the categorization and claims are accurate, the survey would serve as a useful entry point for researchers and practitioners seeking a structured overview of RL applications in manufacturing, energy, and robotics. Its strengths include broad coverage of the literature, clear taxonomies in Tables I–III, and a substantial reference list spanning both application domains and cross-cutting challenges. However, the value of a survey of this kind rests on the reliability of its categorizations and quantitative statements; for that reason, the internal inconsistency in Table II is a significant obstacle to the paper's central contribution. The paper does not provide a reproducible methodology, but as a survey it can be remedied by correcting the tables and adding a short methodology statement.","major_comments":[{"comment":"The 'Renewable Energy Integration' column lists PPO [29], Batch Constrained Munchausen Deep Q-learning [47], Q-learning [48], and A3C [45] as its RL approaches, but all four references are HVAC-control papers: [29] is whole-building HVAC control and demand response, [47] is safe HVAC control via batch RL, [48] is HVAC operation optimization, and [45] is end-to-end DRL for HVAC in office buildings. The narrative text for renewable energy integration instead cites [4], [31], [41]–[44]. This is a verifiable internal contradiction: the table assigns works to a sub-domain that contradicts the paper's own prose. Because Section II's tables are the main deliverable of the survey, this error calls the reliability of the 'systematic categorization' into question until a full audit is performed.","section":"Section II.B, Table II"},{"comment":"The paper claims a 'comprehensive bibliography' and repeatedly refers to 'representative studies' without stating any systematic search strategy, inclusion criteria, or quality-assessment procedure. This makes the representativeness and comprehensiveness of the selected literature unverifiable. The authors should either add a methodology subsection describing how the literature was collected and filtered, or temper the claims of comprehensiveness and of being the 'first survey' to match the actual selection process.","section":"Section I.B"},{"comment":"The statement that demand-response strategies achieve 'up to 22% energy savings' is presented without attribution to a specific experiment or study among the cited references [29]–[34]. In a survey, quantitative performance claims must be traceable to the original source so readers can assess the context, methodology, and generality of the result. Please cite the specific reference(s) and describe the conditions under which this figure was obtained.","section":"Section II.B, Demand Response paragraph"}],"minor_comments":[{"comment":"The phrase 'a effective framework' should be 'an effective framework'.","section":"Section I.A"},{"comment":"The term 'HV AC' should be consistently written as 'HVAC' without a space.","section":"Section II.B and Table II"},{"comment":"Reference [84] is by Zanon and Gros, not by Li et al.; the author attribution in the 'Related Studies' column should be corrected.","section":"Table IV, Safety and Robustness row"},{"comment":"The RL Approaches column lists 'EfficientLPT [52]', which is not an RL algorithm itself but a method that uses prior policy guidance; consider moving it to the methodology highlights or clarifying its role.","section":"Table III, Motion Planning row"},{"comment":"The RL Approaches column lists PPO, SAC, MBPO, Dreamer, IMPALA, and Acme without citations; adding references for these algorithms would improve the completeness of the table.","section":"Table IV, Sample Efficiency and Scalability row"},{"comment":"The footnote states both that 'This work is not supported by any organization' and that the work is a preprint of a paper published at IEEE CASE 2024; please clarify the relationship between the two versions and any copyright or overlap considerations.","section":"Title footnote"}],"recommendation":"major_revision","confidential_remarks":"The Table II mismatch is serious enough that I recommend a full audit of all three application tables before acceptance. A spot-check of Tables I and III did not reveal similar sub-domain misassignments, but the audit should be systematic and the authors should document the process. The lack of a methodology section is a secondary concern that should be addressed in the same revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a genuinely useful survey for someone entering RL-for-automation, but its main deliverable—the application tables—has a verifiable error. In Table II, the Renewable Energy Integration row lists PPO [29], Batch Constrained Munchausen Deep Q-learning [47], Q-learning [48], and A3C [45]. Checking the references shows all four are HVAC-control papers: whole-building HVAC/demand response, safe HVAC batch RL, HVAC operation optimization, and end-to-end DRL for HVAC in office buildings. None is about renewable integration. The surrounding prose for renewable integration cites [4], [31], [41]-[44], which are the right kind of references. So the paper contradicts itself, and the table—the thing a newcomer will use—misleads.\n\nWhat's good: the survey covers manufacturing, energy, and robotics with a consistent template, gives algorithm names and representative studies, and has a solid challenges section (sample efficiency, safety, interpretability, transfer, deployment) with a useful table. The bibliography is genuinely broad and current enough for a 2024 CASE paper. For someone scoping a literature search, it saves time.\n\nSoft spots, in proportion. The stress-test issue above is real and should be fixed. Second, the 'representative studies' are selected with no stated criteria, so the bounds of the survey are unverifiable; a reader can't tell if omissions are deliberate or accidental. Third, the 'first survey specifically focused on RL for optimization in automation' claim is plausible but unproven—they don't compare against existing surveys to establish the gap. There are also smaller slips ('a effective', a stray parenthesis in the reference list) that a careful proofread would catch.\n\nI don't see a load-bearing methodological flaw beyond the table error. The taxonomy is reasonable, and the prose is mostly consistent with the cited literature. But the table error is not minor: Table II is the centerpiece of the energy section, and it will mislead. This deserves peer review, because the fix is straightforward and the survey has real value. I would not cite it in my own work until the table is corrected.","headline":"A useful organizing survey for newcomers, but Table II misassigns four HVAC papers to renewable energy integration, so the core categorization needs a fix before it can be trusted.","tokens_in":14516,"tokens_out":2033,"would_cite":false,"duration_ms":18122,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that reinforcement learning has become a practical tool for optimization across manufacturing, energy systems, and robotics, and that it provides the first structured map of the field, organizing representative studies…","keywords":["reinforcement learning","automation","manufacturing","energy systems","robotics","sample efficiency","safety","interpretability"],"falsifier":"A reader could test the paper's comprehensiveness by running a systematic literature search for RL optimization studies in each of the twelve subdomains over the same period, using explicit inclusion criteria, and checking whether the survey's cited references cover, say, the majority of qualifying papers. If large numbers of qualifying papers are missing, the survey's claim to be a comprehensive guide would be refuted; if its cites dominate the qualifying set, the claim stands.","tokens_in":13631,"feed_emoji":"🤖","tokens_out":4525,"duration_ms":36723,"temperature":0.7,"pith_summary":"The paper argues that reinforcement learning has become a practical tool for optimization problems in automation and that the field has matured enough to be surveyed. It sets out to organize the literature into three application domains—manufacturing, energy systems, and robotics—and into five cross-cutting challenge areas: sample efficiency and scalability, safety and robustness, interpretability and trustworthiness, transfer learning and meta-learning, and real-world deployment and integration. The authors claim this is the first survey devoted specifically to RL for optimization in automation, and they provide comparative tables that classify representative studies by objectives, RL methods, outcomes, and future directions. A sympathetic reader would take the paper's contribution to be a structured map of the field that helps researchers locate relevant work and open problems.","feed_headline":"First survey maps RL optimization in manufacturing, energy, robotics","feed_subtitle":"Categorizes methods, challenges, and open problems across three automation domains to guide new work.","key_machinery":"The central organizing device is a set of comparison tables (Tables I–IV) that act as a taxonomy. Each table rows subdomains against feature columns—key objectives, challenges addressed, RL approaches, methodology highlights, outcomes, future directions, and representative studies. Table IV adds a cross-cutting view of five challenge areas, pairing each with its state of the art and future directions. These tables carry the argument: they are how the survey converts scattered individual papers into a structured claim about the field's state and needs.","core_discovery":"On the authors' own terms, the discovery is that RL-based optimization in automation has reached a state where it can be systematically categorized, and that the field's progress clusters around identifiable methods (DQN, PPO, SAC, DDPG, MARL variants) applied to recurring problem types (scheduling, inventory, maintenance, process control, demand response, microgrid management, renewable integration, HVAC control, motion planning, manipulation, multi-robot coordination, human-robot collaboration). The survey's central claim is that this categorization is the first of its kind specifically for RL in automation, that it reveals common methodological patterns across domains, and that five shared challenges—sample efficiency, safety, interpretability, transfer, deployment—currently limit real-world use. It also claims to identify the state of the art within each subdomain and to list future research directions that follow from the gaps it finds.","pith_inferences":["The survey's definition of 'optimization' is broad enough to cover almost any RL application in these domains, so its claim of being the first such survey depends on how tightly 'optimization' is scoped; a narrower reading might find earlier focused reviews.","The emphasis on representative studies rather than an exhaustive census suggests that the survey's tables are best read as a starting point, not a complete map; a systematic, reproducible search with inclusion criteria could extend or correct the coverage.","Since the paper points to transfer learning and meta-learning as future directions, one testable extension is to benchmark whether cross-domain RL policies trained in manufacturing scheduling can transfer to energy dispatch or robot planning tasks with minimal fine-tuning.","The repeated mention of safety and trust implies that adoption of RL in automation will depend on progress in formal verification, not just algorithmic performance; a survey tracking verification methods alongside RL would be a natural sequel."],"forward_implications":["Researchers entering any of the twelve subdomains can use the tables to find representative baseline works and standard RL algorithms without a separate literature search.","The consistent pattern of challenges across domains implies that advances in sample efficiency, safety, or interpretability in one domain should transfer in method to the others.","The gap between simulated successes and real-world deployment, identified as a cross-cutting challenge, suggests that future work on sim-to-real transfer and human-in-the-loop integration would have broad impact.","The survey's categorization implies that multi-agent RL, currently prominent in inventory, microgrid, and multi-robot settings, is a general template for distributed automation optimization.","The identified future directions (risk-sensitive formulations, curriculum learning, explainable RL, meta-RL) provide a shared research agenda for the field."],"supporting_citations":[{"why":"Supplies the foundational formalization of RL as sequential decision-making that the survey uses to frame all applications.","marker":"[1]"},{"why":"Establishes deep RL's breakthrough capability, the basis for the DQN-family methods recurring across the survey's tables.","marker":"[2]"},{"why":"Prior review of deep RL in smart manufacturing that the survey positions itself against and extends with cross-domain coverage.","marker":"[3]"},{"why":"Prior review of RL in energy systems that grounds the energy-table's categorization.","marker":"[4]"},{"why":"Prior survey of RL in robotics that grounds the robotics applications section.","marker":"[5]"}],"fun_headline_variants":["First RL optimization survey covers manufacturing, energy, robotics","Survey reveals how RL optimizes automation across three sectors","RL for automation optimization: first survey maps methods and gaps","Survey identifies five challenges in RL-driven automation optimization","First taxonomy of RL optimization for manufacturing, energy, robots"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey assumes that the handful of 'representative studies' it selects for each subdomain accurately reflects the state of the art, without describing a systematic search strategy or inclusion criteria.","fun_headline_variants_meta":{"raw":{"variants":["First RL optimization survey covers manufacturing, energy, robotics","Survey reveals how RL optimizes automation across three sectors","RL for automation optimization: first survey maps methods and gaps","Survey identifies five challenges in RL-driven automation optimization","First taxonomy of RL optimization for manufacturing, energy, robots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000911,"raw_usage":{"total_tokens":3875,"prompt_tokens":866,"completion_tokens":3009,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":482,"completion_tokens_details":{"reasoning_tokens":2932}},"tokens_in":482,"tokens_out":3009,"duration_ms":18943,"temperature":1.0,"reasoning_tokens":2932,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T21:33:45.229110+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could test the paper's comprehensiveness by running a systematic literature search for RL optimization studies in each of the twelve subdomains over the same period, using explicit inclusion criteria, and checking whether the survey's cited references cover, say, the majority of qualifying papers. If large numbers of qualifying papers are missing, the survey's claim to be a comprehensive guide would be refuted; if its cites dominate the qualifying set, the claim stands.","supporting_citations":[],"review_version":1}