REVIEW 4 major objections 4 minor 54 references
CIRO7.2: A Material Network with Circularity of -7.2 and Reinforcement-Learning-Controlled Robotic Disassembler
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A sensitivity law ties robot disassembly success to circularity: each failed task costs more when parts are heavier or scarcer.
desk verdict A coherent metric extension and demonstration, but the headline sensitivity 'finding' is a definitional artifact and the RL evidence is too thin to carry it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the time-window circularity $\lambda(\mathcal{N};t)$, defined as the negative time-average of finite-time-sustainable material leaving the network, weighted by a criticality coefficient and by a functionality factor that penalizes discarding working goods. The network is a directed graph of thermodynamic compartments, and the robot is one such compartment whose policy produces the success fraction $s$ and disassembly time $T_d$; mass conservation splits the incoming criticality-weighted mass $m_0$ into a reused part $m_r=m_0 s/100$ and an incinerated part $m_u=m_0(1-s/100)$. The analytical engine is Eq. (16), in which circularity factorizes as $-2m_0\,\alpha(s)$ with $\alpha(s)$ independent of $m_0$; that factorization is what converts the simulation results into a sensitivity claim about material mass and criticality.
What would settle it
Measure the four time scales for a real disassembled product, evaluate Eq. (12) directly with the exact final time $t_f$, and check two consequences of the paper's claim: the $\lambda$ ordering across the four tasks should follow Eq. (20), and the sensitivity of $\lambda$ to $s$ should increase with the criticality-weighted mass $m_0$. A dataset where either consequence fails would falsify the sensitivity law outside the assumed time regime.
Extended reading notes
Core claim
The paper's central claim is that the quality of an RL controller at a disassembly station is a whole-network quantity: a change in the success fraction $s$ moves $\lambda$ by an amount proportional to $m_0$, the sum of criticality-weighted masses entering the chain. Under the approximation that product-use time and reuse time dominate disassembly, transport, and incineration times, the circularity of this network is $\lambda \approx -\frac{2m_0}{t_{2,\mathrm{in},4}+T_r}(t_{2,\mathrm{in},4}+(2-\frac{s}{100})T_r)$, with limiting cases $\lambda\approx -2m_0$ for perfect disassembly and $\lambda\approx -2m_0\frac{t_{2,\mathrm{in},4}+2T_r}{t_{2,\mathrm{in},4}+T_r}$ for complete failure. The simulated tasks put these limits at $-2.1$ for two 1 kg parts with full success and $-7.2$ for four 1 kg parts inside a 3 kg chassis with no successful disassembly. The paper reads this as a sensitivity theorem: improving the robot policy matters most exactly where the materials are largest and most critical.
Load-bearing premise
The load-bearing premise is that the arbitrary time and criticality parameters in Table I are representative and that reuse time dominates disassembly, transport, and incineration; if that premise fails, the approximate $\lambda$ values change even though the sign of the sensitivity to success may survive.
Editorial extensions
If this is right
- If Eq. (20) holds, the same percentage-point drop in disassembly success costs twice as much circularity on a flow with twice the criticality-weighted mass, so robot performance matters more for heavier and scarcer products.
- At perfect disassembly ($s=100\%$), $\lambda$ becomes approximately $-2m_0$ and barely depends on time, so the main circularity levers left are mass reduction and material substitution.
- For failed disassembly, $\lambda$ approaches $-2m_0$ times a factor between 1 and 2 that depends on how long reused material stays in circulation, making reuse time an explicit second-order knob.
- The four simulated tasks show the predicted monotone drop: $\lambda$ goes from $-2.1$ for two 1 kg parts with full success to $-3.1$ and $-6.3$ for failed tasks with more material, reaching $-7.2$ when a 3 kg chassis is added.
Reading between the lines
- Not stated in the paper: because Eq. (16) is linear in $s$, episode-to-episode variance in disassembly success should not affect $\lambda$; only the mean success rate matters, so two policies with the same average success should give the same circularity.
- Not stated in the paper: across a mixed portfolio of end-of-use products, the formula suggests prioritizing reinforcement-learning effort toward the flow with the largest $c_{f,b,i} m^0_{f,b,i}$, since that flow dominates the circularity penalty.
- Not stated in the paper: replacing the Table I times with measured supply-chain durations for a specific product category would turn the approximate law into a testable engineering prediction, and the paper reports no such measurement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a thermodynamical material network that includes a reinforcement-learning-controlled robotic disassembler, and defines a time-window circularity metric λ. Four simulated disassembly tasks of increasing complexity are trained with SAC, TQC, and TD3 (each enhanced with HER) in the panda-gym simulator, and the resulting circularity values are reported (from −2.1 to −7.2). The authors derive an approximate expression for λ in Eq. (20) and claim that the impact of RL controller performance on circularity has a positive correlation with the mass and criticality of the disassembled materials. The paper also proposes 'circular intelligence and robotics' (CIRO) as an emerging research field.
Significance. The analytic derivation of λ in Eq. (12) and its sensitivity in Eq. (20) is transparent and correct, and the source code is publicly available. However, the central 'positive correlation' claim is a direct algebraic consequence of the definition of λ, since λ is linear in the weighted mass m0 and the factor α(s) is independent of m0; thus it is not an empirical result established by the RL experiments. The RL experiments are preliminary: each configuration is trained once, and the more complex tasks are deliberately trained for far fewer steps than the simple task, with all controllers achieving s=0 in those cases. The headline value λ=−7.2 is also affected by an internal inconsistency in the chassis mass text. The paper is valuable as an illustrative, mathematically explicit case study, but its claims need to be reframed and its reporting tightened.
major comments (4)
- [Section III-C and Section III-D] The positive-correlation claim (abstract, Section III-C after Table IV, Section III-D) is an algebraic consequence of the metric definition, not an empirically validated finding. Because λ in Eq. (12) is linear in m0 and α(s) in Eq. (17) is independent of m0, Eq. (20) shows by construction that the sensitivity of λ to s scales with m0. The RL experiments do not test this correlation: in Tables III, IV, and V all controllers yield s=0, so no variation in s is observed across different masses or criticalities. The paper should state that this is an analytic property of the metric, and should avoid presenting it as if it were established by the simulations.
- [Section III-C, chassis task paragraph] The chassis scenario contains an internal inconsistency: the text states m0_f,b,1=2 kg and m0_f,b,2=5 kg, which gives m0=0.1×2+0.95×5=4.95 kg, yet the paragraph and Table V report m0=2.4 kg. The table values (5 and 2) are the only ones consistent with 2.4 kg. This error directly affects the reproducibility of the headline value λ=−7.2 and must be corrected in the text.
- [Section III-C, Tables III and IV] The conclusion that more complex tasks are harder for RL is under-supported because the harder tasks are deliberately trained for fewer steps than Task 1 (1.5×10^5 and 2.0×10^5 vs. 4.5×10^5) and all algorithms achieve s=0 in those tasks. Additionally, every configuration is run only once, so no variance across random seeds is reported. To support the ordering of task difficulty, the paper should either train the harder tasks to convergence or provide clear evidence of convergence (or saturation), and should report multiple seeds with statistics.
- [Section III-D and Table I] The numerical values of λ and the sensitivity result depend on arbitrary parameter choices (t2,in,4=1 month, Tr=1 month, Ti=1 day, Tt=1 hour) and on the approximation tf≈t2,in,4+Tr in Eq. (15). The paper should state explicitly that these parameters are illustrative and not calibrated to a real supply chain, and should discuss how the results change when Eq. (15) is not valid (e.g., when product use time is not dominant).
minor comments (4)
- [Section III-D, Eq. (17)] In the text after Eq. (17), the statement 'with α(s)∈[1,2) since s∈[0,1]' is inconsistent with the earlier definition of s∈[0,100] in Section III-B; it should read 'since s/100∈[0,1]'.
- [Table II, TD3 row] The Td value for TD3-HER is reported as '186400*' but the footnote says Td=86400 seconds; this appears to be a typo for '86400*'.
- [Abstract and Fig. 1 caption] The phrase 'An higher circularity' should be 'A higher circularity'; the abstract also lacks a comma after 'take-make-dispose' in the first sentence.
- [Section III-C, first paragraph] The claim that one simulated time step equals approximately 40 ms of real time is taken from panda-gym but no justification or reference is given; a brief explanation would help the reader assess the Td values.
Circularity Check
The paper's central sensitivity finding is an algebraic restatement of the λ definition, not an empirical result; the RL experiments cannot test it because all high-mass tasks yield s=0.
-
self definitional
[Section III-D, after Eq. (20)]
"sinceα is independent of m0, the sensitivity ofλtoshas a positive correlation with m0, which confirms that the importance of a performing RL controller increases when there is an increase of the masses and/or criticality of the materials to be disassembled"
From Eq. (12), λ is a linear function of m0 and of the criticality coefficients c_f,b,i entering m0 via Eq. (6). Under the approximation (15), Eq. (20) gives λ ≈ -2m0 α(s) with α(s) independent of m0. Therefore the magnitude of ∂λ/∂s equals 2m0 |α'(s)|, which is proportional to m0 by construction. The claimed positive correlation of RL-performance impact with material quantity and criticality is thus a direct consequence of how λ is defined, not a finding established by the experiments. The RL data do not provide independent support: in the higher-mass tasks (Tables IV and V) all controllers yield s=0, so no variation of s is observed, and the 'increasing impact' is read off the formula rather than measured.
full rationale
The central sensitivity claim is forced by the definition of λ: Eq. (12) makes λ linear in m0 = c_f,b,1 m0_f,b,1 + c_f,b,2 m0_f,b,2, and Eq. (20) reduces the model to λ ≈ -2m0 α(s) with α(s)∈[1,2). Hence the derivative of λ with respect to s is proportional to m0 by construction, making the 'positive correlation with quantity and criticality' an algebraic restatement rather than an empirical discovery. The RL experiments do not test this correlation because in all tasks with larger m0 (Tables III, IV, V) every controller gives s=0; the only nontrivial s variation occurs in the 2-part/1-target task. The absolute λ values (-2.1, -3.1, -6.3, -7.2) are direct evaluations of Eq. (12) using the hand-chosen Table I parameters, so they are not fitted, but they inherit the definition and are not independently validated. There is also a reproducibility inconsistency in the chassis task: the text states m0_f,b,1=2 kg and m0_f,b,2=5 kg and then reports m0=2.4 kg, whereas the table values (5,2) are the ones that yield 0.1*5+0.95*2=2.4 kg; this is a correctness issue, not a circularity issue. The paper does rely on the authors' prior TMN/λ work (Refs. [17], [24], [25]), but it restates the definition and performs the algebra itself, so the self-citation is not load-bearing. Overall, the paper is self-contained, but its headline finding is forced by the metric's definition, giving a circularity score of 8.
Assumptions & free parameters
free parameters (9)
- l (functionality coefficient) =
1
- c_f,b,1 (criticality of beta1) =
0.1
- c_f,b,2 (criticality of beta2) =
0.95
- t2,in,4 (time from extraction to disassembly facility) =
2,592,000 s (1 month)
- Tr (reuse time) =
2,592,000 s (1 month)
- Ti (incineration time) =
86,400 s (1 day)
- Tt (transport time to incinerator) =
3,600 s (1 hour)
- Delta (time conversion constant in (2)) =
unspecified, must be fixed across calculations
- Simulation step to real time conversion =
40 ms per step
assumptions (6)
- domain assumption Compartmental dynamical thermodynamics framework (Haddad 2019)
- domain assumption Mass conservation m0 = mu + mr
- ad hoc to paper Definition of finite-time sustainable mass (Definition 3)
- ad hoc to paper Non-disassembled parts are considered functional (l=1)
- domain assumption Approximation tf ≈ t2,in,4 + Tr
- domain assumption All parts weigh 1 kg and are made of beta1 or beta2
invented entities (1)
-
Circular intelligence and robotics (CIRO)
Cite this review
Pith. "Pith review of CIRO7.2: A Material Network with Circularity of -7.2 and Reinforcement-Learning-Controlled Robotic Disassembler." pith.science (2026). https://pith.science/paper/TLMJX7LE
@misc{pith2026250611748,
author = {Pith},
title = {Pith review of: CIRO7.2: A Material Network with Circularity of -7.2 and Reinforcement-Learning-Controlled Robotic Disassembler},
year = {2026},
howpublished = {\url{https://pith.science/paper/TLMJX7LE}},
note = {Machine review of arXiv:2506.11748}
}
abstract
The competition over natural reserves of minerals is expected to increase in part because of the linear-economy paradigm based on take-make-dispose. Simultaneously, the linear economy considers end-of-use products as waste rather than as a resource, which results in large volumes of waste whose management remains an unsolved problem. Since a transition to a circular economy can mitigate these open issues, in this paper we begin by enhancing the notion of circularity based on compartmental dynamical thermodynamics, namely, $\lambda$, and then, we model a thermodynamical material network processing a batch of 2 solid materials of criticality coefficients of 0.1 and 0.95, with a robotic disassembler compartment controlled via reinforcement learning (RL), and processing 2-7 kg of materials. Subsequently, we focused on the design of the robotic disassembler compartment using state-of-the-art RL algorithms and assessing the algorithm performance with respect to $\lambda$ (Fig. 1). The highest circularity is -2.1 achieved in the case of disassembling 2 parts of 1 kg each, whereas it reduces to -7.2 in the case of disassembling 4 parts of 1 kg each contained inside a chassis of 3 kg. Finally, a sensitivity analysis highlighted that the impact on $\lambda$ of the performance of an RL controller has a positive correlation with the quantity and the criticality of the materials to be disassembled. This work also gives the principles of the emerging research fields indicated as circular intelligence and robotics (CIRO). Source code is publicly available.
Figures
Reference graph
Works this paper leans on
-
[17]
Circular Economy Design through System Dynamics Modeling
F. Zocco and M. Malvezzi, “Circular economy design through system dynamics modeling,”arXiv preprint arXiv:2411.13540, 2024
work page Pith review arXiv 2024
-
[1]
CRM Alliance, 2024, webpage: https://www.crmalliance.eu/ critical-raw-materials; last access: 28 April 2025
work page 2024
-
[2]
European Commission, “Critical raw materials,” 2023, webpage: https://single-market-economy.ec.europa.eu/sectors/raw-materials/ areas-specific-interest/critical-raw-materials en; last access: 28 April 2025
work page 2023
-
[3]
Notice of final determination on 2023 DOE critical materials list,
Federal Register, “Notice of final determination on 2023 DOE critical materials list,” 2023, webpage: https: //www.federalregister.gov/documents/2023/08/04/2023-16611/ notice-of-final-determination-on-2023-doe-critical-materials-list; last access: 28 April 2025
work page 2023
-
[4]
eurostat, “Waste statistics,” 2024, available at: https://ec.europa. eu/eurostat/statistics-explained/index.php?title=Waste statistics; last ac- cess: 28 April 2025
work page 2024
-
[5]
Critical minerals and great power com- petition: An overview,
J. Zhou and A. M ˚anberger, “Critical minerals and great power com- petition: An overview,”STOCKHOLM INTERNATIONAL PEACE RE- SEARCH INSTITUTE, October 2024
work page 2024
-
[6]
A. Nygaard, “The geopolitical risk and strategic uncertainty of green growth after the Ukraine invasion: How the circular economy can decrease the market power of and resource dependency on critical minerals,”Circular Economy and Sustainability, vol. 3, no. 2, pp. 1099– 1126, 2023
work page 2023
-
[7]
Recycling and circular econ- omy—Towards a closed loop for metals in emerging clean technologies,
C. Hagel ¨uken and D. Goldmann, “Recycling and circular econ- omy—Towards a closed loop for metals in emerging clean technologies,” Mineral Economics, vol. 35, no. 3, pp. 539–562, 2022
work page 2022
Show all 54 references
-
[8]
Towards circular economy implementation: a comprehensive review in context of manufacturing industry,
M. Lieder and A. Rashid, “Towards circular economy implementation: a comprehensive review in context of manufacturing industry,”Journal of Cleaner Production, vol. 115, pp. 36–51, 2016
2016
-
[9]
Demystifying series: Business and a circular economy,
CE-Hub Team, “Demystifying series: Business and a circular economy,” 2022, available at: https://exeterce.org/knowledge-hub/ demystifying-business-and-a-circular-economy/# ftnref4; last access: 28 April 2025
2022
-
[10]
Metal scarcity and sustain- ability, analyzing the necessity to reduce the extraction of scarce metals,
M. Henckens, P. Driessen, and E. Worrell, “Metal scarcity and sustain- ability, analyzing the necessity to reduce the extraction of scarce metals,” Resources, Conservation and recycling, vol. 93, pp. 1–8, 2014
2014
-
[11]
Scarce mineral resources: Extraction, consumption and limits of sustainability,
T. Henckens, “Scarce mineral resources: Extraction, consumption and limits of sustainability,”Resources, Conservation and Recycling, vol. 169, p. 105511, 2021
2021
-
[12]
Sustainability,
United Nations, “Sustainability,” available at: https://www.un.org/en/ academic-impact/sustainability; last access: 28 April 2025
2025
-
[13]
What is a circular economy?
Ellen MacArthur Foundation, “What is a circular economy?” webpage: https://www.ellenmacarthurfoundation.org/topics/ circular-economy-introduction/overview; last access: 28 April 2025
2025
-
[14]
Digitalisation for the transition to a resource efficient and circular economy,
E. Bartekov ´a and P. B ¨orkey, “Digitalisation for the transition to a resource efficient and circular economy,”OECD Environment Working Papers, No. 192, OECD Publishing, Paris, 2022, https://doi.org/10.1787/ 6f6d18e7-en
2022
-
[15]
Towards a thermodynamical deep-learning-vision-based flexible robotic cell for circular healthcare,
F. Zocco, D. Sleath, and S. Rahimifard, “Towards a thermodynamical deep-learning-vision-based flexible robotic cell for circular healthcare,” Circular Economy and Sustainability, pp. 1–23, 2025
2025
-
[16]
Introduction to an adaptive strategy for circular design,
Ellen MacArthur Foundation, “Introduction to an adaptive strategy for circular design,” webpage: https://www.ellenmacarthurfoundation.org/ adaptive-strategy-for-circular-design/introduction; last access: 28 April 2025
2025
-
[18]
W. M. Haddad,A Dynamical Systems Theory of Thermodynamics. Princeton University Press, 2019
2019
-
[19]
Circular economy indicators for supply chains: A systematic literature review,
T. Calzolari, A. Genovese, and A. Brint, “Circular economy indicators for supply chains: A systematic literature review,”Environmental and Sustainability Indicators, vol. 13, p. 100160, 2022
2022
-
[20]
The future of circular economy metrics: Expert visions,
M. Saidani, T. Shevchenko, Z. S. Esfandabadi, M. Ranjbari, J. A. Mesa, B. Yannou, and F. Cluzel, “The future of circular economy metrics: Expert visions,”Resources, Conservation and Recycling, vol. 205, p. 107565, 2024
2024
-
[21]
The hidden concept and the beauty of multiple “R
A. A. Zorpas, “The hidden concept and the beauty of multiple “R” in the framework of waste strategies development reflecting to circular economy principles,”Science of the Total Environment, vol. 952, p. 175508, 2024
2024
-
[22]
Circular economy: Measuring innovation in the product chain,
J. Potting, M. P. Hekkert, E. Worrell, and A. Hanemaaijer, “Circular economy: Measuring innovation in the product chain,”Planbureau voor de Leefomgeving, no. issue 2544, report, 2017
2017
-
[23]
Developing a strategic methodology for circular economy roadmapping: A theoretical framework,
H. Abu-Bakar and F. Charnley, “Developing a strategic methodology for circular economy roadmapping: A theoretical framework,”Sustainabil- ity, vol. 16, no. 15, p. 6682, 2024
2024
-
[24]
Thermodynami- cal material networks for modeling, planning, and control of circular ma- terial flows,
F. Zocco, P. Sopasakis, B. Smyth, and W. M. Haddad, “Thermodynami- cal material networks for modeling, planning, and control of circular ma- terial flows,”International Journal of Sustainable Engineering, vol. 16, no. 1, pp. 1–14, 2023
2023
-
[25]
Circularity of thermodynamical material networks: Indicators, examples, and algorithms,
F. Zocco, “Circularity of thermodynamical material networks: Indicators, examples, and algorithms,”arXiv preprint arXiv:2209.15051, 2024
2024
-
[26]
Synchronized object detection for autonomous sorting, mapping, and quantification of JOURNAL OF LATEX CLASS FILES, VOL. 00, NO. 0, MONTH 20XX 8 materials in circular healthcare,
F. Zocco, D. R. Lake, S. McLoone, and S. Rahimifard, “Synchronized object detection for autonomous sorting, mapping, and quantification of JOURNAL OF LATEX CLASS FILES, VOL. 00, NO. 0, MONTH 20XX 8 materials in circular healthcare,”arXiv preprint arXiv:2405.06821, 2024, to app...
2024 arXiv
-
[27]
Material flows and efficiency,
J. M. Cullen and D. R. Cooper, “Material flows and efficiency,”Annual Review of Materials Research, vol. 52, no. 1, pp. 525–559, 2022
2022
-
[28]
P. H. Brunner and H. Rechberger,Handbook of material flow analysis: For environmental, resource, and waste engineers. CRC press, 2016
2016
-
[29]
A unification be- tween deep-learning vision, compartmental dynamical thermodynamics, and robotic manipulation for a circular economy,
F. Zocco, W. M. Haddad, A. Corti, and M. Malvezzi, “A unification be- tween deep-learning vision, compartmental dynamical thermodynamics, and robotic manipulation for a circular economy,”IEEE Access, vol. 12, pp. 173 502–173 516, 2024
2024
-
[30]
Robot for automatic waste sorting on construction sites,
X. Chen, H. Huang, Y . Liu, J. Li, and M. Liu, “Robot for automatic waste sorting on construction sites,”Automation in Construction, vol. 141, p. 104387, 2022
2022
-
[31]
Robotic waste sorting technology: Toward a vision- based categorization system for the industrial robotic separation of recyclable waste,
M. Koskinopoulou, F. Raptopoulos, G. Papadopoulos, N. Mavrakis, and M. Maniadakis, “Robotic waste sorting technology: Toward a vision- based categorization system for the industrial robotic separation of recyclable waste,”IEEE Robotics & Automation Magazine, vol. 28, no. 2, pp...
2021
-
[32]
Challenges for future robotic sorters of mixed industrial waste: a survey,
T. Kiyokawa, J. Takamatsu, and S. Koyanaka, “Challenges for future robotic sorters of mixed industrial waste: a survey,”IEEE Transactions on Automation Science and Engineering, vol. 21, no. 1, pp. 1023–1040, 2022
2022
-
[33]
Laili, Y
Y . Laili, Y . Wang, Y . Fang, and D. T. Pham,Optimisation of robotic disassembly for remanufacturing. Springer, 2022
2022
-
[34]
Enhancing disassembly practices for electric vehicle battery packs: a narrative comprehensive review,
M. Beghi, F. Braghin, and L. Roveda, “Enhancing disassembly practices for electric vehicle battery packs: a narrative comprehensive review,” Designs, vol. 7, no. 5, p. 109, 2023
2023
-
[35]
Robotic disassembly for increased recovery of strategically important materials from electrical vehicles,
J. Li, M. Barwood, and S. Rahimifard, “Robotic disassembly for increased recovery of strategically important materials from electrical vehicles,”Robotics and Computer-Integrated Manufacturing, vol. 50, pp. 203–212, 2018. [Online]. Available: https://www.sciencedirect.com/ scie...
2018
-
[36]
Robotic disassembly task training and skill transfer using reinforcement learning,
M. Qu, Y . Wang, and D. T. Pham, “Robotic disassembly task training and skill transfer using reinforcement learning,”IEEE Transactions on Industrial Informatics, vol. 19, no. 11, pp. 10 934–10 943, 2023
2023
-
[37]
Human- robot collaborative disassembly in industry 5.0: A systematic literature review and future research agenda,
G. Yuan, X. Liu, X. Qiu, P. Zheng, D. T. Pham, and M. Su, “Human- robot collaborative disassembly in industry 5.0: A systematic literature review and future research agenda,”Journal of Manufacturing Systems, vol. 79, pp. 199–216, 2025
2025
-
[38]
Transferring policy of deep reinforcement learning from simulation to reality for robotics,
H. Ju, R. Juan, R. Gomez, K. Nakamura, and G. Li, “Transferring policy of deep reinforcement learning from simulation to reality for robotics,” Nature Machine Intelligence, vol. 4, no. 12, pp. 1077–1087, 2022
2022
-
[39]
How to train your robot with deep reinforcement learning: lessons we have learned,
J. Ibarz, J. Tan, C. Finn, M. Kalakrishnan, P. Pastor, and S. Levine, “How to train your robot with deep reinforcement learning: lessons we have learned,”The International Journal of Robotics Research, vol. 40, no. 4-5, pp. 698–721, 2021
2021
-
[40]
A coach-based bayesian reinforcement learning method for snake robot control,
Y . Jia and S. Ma, “A coach-based bayesian reinforcement learning method for snake robot control,”IEEE Robotics and Automation Letters, vol. 6, no. 2, pp. 2319–2326, 2021
2021
-
[41]
Reinforcement learning models and algorithms for diabetes management,
K.-L. A. Yau, Y .-W. Chong, X. Fan, C. Wu, Y . Saleem, and P.- C. Lim, “Reinforcement learning models and algorithms for diabetes management,”IEEE Access, vol. 11, pp. 28 391–28 415, 2023
2023
-
[42]
Towards a multi-agent reinforcement learning approach for joint sensing and sharing in cognitive radio networks,
K. Rapetswa and L. Cheng, “Towards a multi-agent reinforcement learning approach for joint sensing and sharing in cognitive radio networks,”Intelligent and Converged Networks, vol. 4, no. 1, pp. 50–75, 2023
2023
-
[43]
Tc-driver: A trajectory conditioned reinforcement learning approach to zero-shot autonomous racing,
E. Ghignone, N. Baumann, and M. Magno, “Tc-driver: A trajectory conditioned reinforcement learning approach to zero-shot autonomous racing,”Field Robotics, vol. 3, pp. 637–651, 2023
2023
-
[44]
Reinforcement learning in robotic applications: a comprehensive survey,
B. Singh, R. Kumar, and V . P. Singh, “Reinforcement learning in robotic applications: a comprehensive survey,”Artificial Intelligence Review, vol. 55, no. 2, pp. 945–990, 2022
2022
-
[45]
J. A. Bondy and U. S. R. Murty,Graph Theory with Applications. Macmillan London, 1976, vol. 290
1976
-
[46]
M. J. Moran, H. N. Shapiro, D. D. Boettner, and M. B. Bailey, Fundamentals of Engineering Thermodynamics. John Wiley & Sons, 2010
2010
-
[47]
Thermodynamics: The unique universal science,
W. M. Haddad, “Thermodynamics: The unique universal science,” Entropy, vol. 19, no. 11, p. 621, 2017
2017
-
[48]
Circulate products and materials,
Ellen MacArthur Foundation, “Circulate products and materials,” 2022, webpage: https://www.ellenmacarthurfoundation.org/ circulate-products-and-materials; last access: 17 April 2025
2022
-
[49]
panda-gym: Open-source goal-conditioned environments for robotic learning,
Q. Gallou ´edec, N. Cazin, E. Dellandr ´ea, and L. Chen, “panda-gym: Open-source goal-conditioned environments for robotic learning,” in 4th Robot Learning Workshop: Self-Supervised and Lifelong Learning@ NeurIPS 2021, 2021
2021
-
[50]
Stable-baselines3: Reliable reinforcement learning implementa- tions,
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dor- mann, “Stable-baselines3: Reliable reinforcement learning implementa- tions,”Journal of Machine Learning Research, vol. 22, no. 268, pp. 1–8, 2021
2021
-
[51]
Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,” inInternational Conference on Machine Learning. Pmlr, 2018, pp. 1861–1870
2018
-
[52]
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics,
A. Kuznetsov, P. Shvechikov, A. Grishin, and D. Vetrov, “Controlling overestimation bias with truncated mixture of continuous distributional quantile critics,” inInternational Conference on Machine Learning. PMLR, 2020, pp. 5556–5566
2020
-
[53]
Addressing function approxi- mation error in actor-critic methods,
S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approxi- mation error in actor-critic methods,” inInternational Conference on Machine Learning. PMLR, 2018, pp. 1587–1596
2018
-
[54]
Hindsight experience replay,
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba, “Hindsight experience replay,”Advances in Neural Information Processing Systems, vol. 30, 2017
2017
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.