Pith. sign in

REVIEW 4 major objections 4 minor 54 references

CIRO7.2: A Material Network with Circularity of -7.2 and Reinforcement-Learning-Controlled Robotic Disassembler

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A sensitivity law ties robot disassembly success to circularity: each failed task costs more when parts are heavier or scarcer.

desk verdict A coherent metric extension and demonstration, but the headline sensitivity 'finding' is a definitional artifact and the RL evidence is too thin to carry it. read the letter →

arxiv 2506.11748 v1 pith:TLMJX7LE submitted 2025-06-13 cs.RO cs.CY

classification cs.ROcs.CY
keywords circulareconomythermodynamicalmaterialnetworkstime-windowcircularityreinforcementlearningroboticdisassemblycriticalitysensitivityanalysisintelligenceandrobotics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to show that a robot learning to disassemble end-of-use products can be evaluated by its effect on the circularity of the whole material chain, not just by task success. It models a six-compartment thermodynamical material network in which material leaves a non-renewable reservoir, is used, passes through a reinforcement-learning-controlled disassembler, and then flows to either reuse or incineration, and it computes the time-window circularity $\lambda$ of that network. The central result is the approximate identity $\lambda \approx -2m_0\alpha(s)$, where $m_0$ is the criticality-weighted mass entering the chain and $\alpha(s)\in[1,2)$ depends on the disassembly success fraction $s$. Because $\alpha(s)$ is independent of $m_0$, the equation says the same drop in disassembly success costs more circularity when the material is heavier or more critical, which is what the simulated tasks confirm with values from $-2.1$ to $-7.2$. If the identity holds, it gives a concrete, physics-grounded reason to spend robotic effort on high-mass, high-criticality product flows.

What carries the argument

The load-bearing object is the time-window circularity $\lambda(\mathcal{N};t)$, defined as the negative time-average of finite-time-sustainable material leaving the network, weighted by a criticality coefficient and by a functionality factor that penalizes discarding working goods. The network is a directed graph of thermodynamic compartments, and the robot is one such compartment whose policy produces the success fraction $s$ and disassembly time $T_d$; mass conservation splits the incoming criticality-weighted mass $m_0$ into a reused part $m_r=m_0 s/100$ and an incinerated part $m_u=m_0(1-s/100)$. The analytical engine is Eq. (16), in which circularity factorizes as $-2m_0\,\alpha(s)$ with $\alpha(s)$ independent of $m_0$; that factorization is what converts the simulation results into a sensitivity claim about material mass and criticality.

What would settle it

Measure the four time scales for a real disassembled product, evaluate Eq. (12) directly with the exact final time $t_f$, and check two consequences of the paper's claim: the $\lambda$ ordering across the four tasks should follow Eq. (20), and the sensitivity of $\lambda$ to $s$ should increase with the criticality-weighted mass $m_0$. A dataset where either consequence fails would falsify the sensitivity law outside the assumed time regime.

Watch

Extended reading notes

Core claim

The paper's central claim is that the quality of an RL controller at a disassembly station is a whole-network quantity: a change in the success fraction $s$ moves $\lambda$ by an amount proportional to $m_0$, the sum of criticality-weighted masses entering the chain. Under the approximation that product-use time and reuse time dominate disassembly, transport, and incineration times, the circularity of this network is $\lambda \approx -\frac{2m_0}{t_{2,\mathrm{in},4}+T_r}(t_{2,\mathrm{in},4}+(2-\frac{s}{100})T_r)$, with limiting cases $\lambda\approx -2m_0$ for perfect disassembly and $\lambda\approx -2m_0\frac{t_{2,\mathrm{in},4}+2T_r}{t_{2,\mathrm{in},4}+T_r}$ for complete failure. The simulated tasks put these limits at $-2.1$ for two 1 kg parts with full success and $-7.2$ for four 1 kg parts inside a 3 kg chassis with no successful disassembly. The paper reads this as a sensitivity theorem: improving the robot policy matters most exactly where the materials are largest and most critical.

Load-bearing premise

The load-bearing premise is that the arbitrary time and criticality parameters in Table I are representative and that reuse time dominates disassembly, transport, and incineration; if that premise fails, the approximate $\lambda$ values change even though the sign of the sensitivity to success may survive.

Editorial extensions

If this is right

  • If Eq. (20) holds, the same percentage-point drop in disassembly success costs twice as much circularity on a flow with twice the criticality-weighted mass, so robot performance matters more for heavier and scarcer products.
  • At perfect disassembly ($s=100\%$), $\lambda$ becomes approximately $-2m_0$ and barely depends on time, so the main circularity levers left are mass reduction and material substitution.
  • For failed disassembly, $\lambda$ approaches $-2m_0$ times a factor between 1 and 2 that depends on how long reused material stays in circulation, making reuse time an explicit second-order knob.
  • The four simulated tasks show the predicted monotone drop: $\lambda$ goes from $-2.1$ for two 1 kg parts with full success to $-3.1$ and $-6.3$ for failed tasks with more material, reaching $-7.2$ when a 3 kg chassis is added.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Not stated in the paper: because Eq. (16) is linear in $s$, episode-to-episode variance in disassembly success should not affect $\lambda$; only the mean success rate matters, so two policies with the same average success should give the same circularity.
  • Not stated in the paper: across a mixed portfolio of end-of-use products, the formula suggests prioritizing reinforcement-learning effort toward the flow with the largest $c_{f,b,i} m^0_{f,b,i}$, since that flow dominates the circularity penalty.
  • Not stated in the paper: replacing the Table I times with measured supply-chain durations for a specific product category would turn the approximate law into a testable engineering prediction, and the paper reports no such measurement.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a thermodynamical material network that includes a reinforcement-learning-controlled robotic disassembler, and defines a time-window circularity metric λ. Four simulated disassembly tasks of increasing complexity are trained with SAC, TQC, and TD3 (each enhanced with HER) in the panda-gym simulator, and the resulting circularity values are reported (from −2.1 to −7.2). The authors derive an approximate expression for λ in Eq. (20) and claim that the impact of RL controller performance on circularity has a positive correlation with the mass and criticality of the disassembled materials. The paper also proposes 'circular intelligence and robotics' (CIRO) as an emerging research field.

Significance. The analytic derivation of λ in Eq. (12) and its sensitivity in Eq. (20) is transparent and correct, and the source code is publicly available. However, the central 'positive correlation' claim is a direct algebraic consequence of the definition of λ, since λ is linear in the weighted mass m0 and the factor α(s) is independent of m0; thus it is not an empirical result established by the RL experiments. The RL experiments are preliminary: each configuration is trained once, and the more complex tasks are deliberately trained for far fewer steps than the simple task, with all controllers achieving s=0 in those cases. The headline value λ=−7.2 is also affected by an internal inconsistency in the chassis mass text. The paper is valuable as an illustrative, mathematically explicit case study, but its claims need to be reframed and its reporting tightened.

major comments (4)
  1. [Section III-C and Section III-D] The positive-correlation claim (abstract, Section III-C after Table IV, Section III-D) is an algebraic consequence of the metric definition, not an empirically validated finding. Because λ in Eq. (12) is linear in m0 and α(s) in Eq. (17) is independent of m0, Eq. (20) shows by construction that the sensitivity of λ to s scales with m0. The RL experiments do not test this correlation: in Tables III, IV, and V all controllers yield s=0, so no variation in s is observed across different masses or criticalities. The paper should state that this is an analytic property of the metric, and should avoid presenting it as if it were established by the simulations.
  2. [Section III-C, chassis task paragraph] The chassis scenario contains an internal inconsistency: the text states m0_f,b,1=2 kg and m0_f,b,2=5 kg, which gives m0=0.1×2+0.95×5=4.95 kg, yet the paragraph and Table V report m0=2.4 kg. The table values (5 and 2) are the only ones consistent with 2.4 kg. This error directly affects the reproducibility of the headline value λ=−7.2 and must be corrected in the text.
  3. [Section III-C, Tables III and IV] The conclusion that more complex tasks are harder for RL is under-supported because the harder tasks are deliberately trained for fewer steps than Task 1 (1.5×10^5 and 2.0×10^5 vs. 4.5×10^5) and all algorithms achieve s=0 in those tasks. Additionally, every configuration is run only once, so no variance across random seeds is reported. To support the ordering of task difficulty, the paper should either train the harder tasks to convergence or provide clear evidence of convergence (or saturation), and should report multiple seeds with statistics.
  4. [Section III-D and Table I] The numerical values of λ and the sensitivity result depend on arbitrary parameter choices (t2,in,4=1 month, Tr=1 month, Ti=1 day, Tt=1 hour) and on the approximation tf≈t2,in,4+Tr in Eq. (15). The paper should state explicitly that these parameters are illustrative and not calibrated to a real supply chain, and should discuss how the results change when Eq. (15) is not valid (e.g., when product use time is not dominant).
minor comments (4)
  1. [Section III-D, Eq. (17)] In the text after Eq. (17), the statement 'with α(s)∈[1,2) since s∈[0,1]' is inconsistent with the earlier definition of s∈[0,100] in Section III-B; it should read 'since s/100∈[0,1]'.
  2. [Table II, TD3 row] The Td value for TD3-HER is reported as '186400*' but the footnote says Td=86400 seconds; this appears to be a typo for '86400*'.
  3. [Abstract and Fig. 1 caption] The phrase 'An higher circularity' should be 'A higher circularity'; the abstract also lacks a comma after 'take-make-dispose' in the first sentence.
  4. [Section III-C, first paragraph] The claim that one simulated time step equals approximately 40 ms of real time is taken from panda-gym but no justification or reference is given; a brief explanation would help the reader assess the Td values.

Circularity Check

1 steps flagged · score 8.0 of 10

The paper's central sensitivity finding is an algebraic restatement of the λ definition, not an empirical result; the RL experiments cannot test it because all high-mass tasks yield s=0.

  1. self definitional [Section III-D, after Eq. (20)]
    "sinceα is independent of m0, the sensitivity ofλtoshas a positive correlation with m0, which confirms that the importance of a performing RL controller increases when there is an increase of the masses and/or criticality of the materials to be disassembled"

    From Eq. (12), λ is a linear function of m0 and of the criticality coefficients c_f,b,i entering m0 via Eq. (6). Under the approximation (15), Eq. (20) gives λ ≈ -2m0 α(s) with α(s) independent of m0. Therefore the magnitude of ∂λ/∂s equals 2m0 |α'(s)|, which is proportional to m0 by construction. The claimed positive correlation of RL-performance impact with material quantity and criticality is thus a direct consequence of how λ is defined, not a finding established by the experiments. The RL data do not provide independent support: in the higher-mass tasks (Tables IV and V) all controllers yield s=0, so no variation of s is observed, and the 'increasing impact' is read off the formula rather than measured.

full rationale

The central sensitivity claim is forced by the definition of λ: Eq. (12) makes λ linear in m0 = c_f,b,1 m0_f,b,1 + c_f,b,2 m0_f,b,2, and Eq. (20) reduces the model to λ ≈ -2m0 α(s) with α(s)∈[1,2). Hence the derivative of λ with respect to s is proportional to m0 by construction, making the 'positive correlation with quantity and criticality' an algebraic restatement rather than an empirical discovery. The RL experiments do not test this correlation because in all tasks with larger m0 (Tables III, IV, V) every controller gives s=0; the only nontrivial s variation occurs in the 2-part/1-target task. The absolute λ values (-2.1, -3.1, -6.3, -7.2) are direct evaluations of Eq. (12) using the hand-chosen Table I parameters, so they are not fitted, but they inherit the definition and are not independently validated. There is also a reproducibility inconsistency in the chassis task: the text states m0_f,b,1=2 kg and m0_f,b,2=5 kg and then reports m0=2.4 kg, whereas the table values (5,2) are the ones that yield 0.1*5+0.95*2=2.4 kg; this is a correctness issue, not a circularity issue. The paper does rely on the authors' prior TMN/λ work (Refs. [17], [24], [25]), but it restates the definition and performs the algebra itself, so the self-citation is not load-bearing. Overall, the paper is self-contained, but its headline finding is forced by the metric's definition, giving a circularity score of 8.

Assumptions & free parameters 9 free parameters · 6 assumptions · 1 invented entities

The λ values are computed from the authors' own definition (2) using the hand-picked parameters above; no free parameter is fitted to external data. The sensitivity analysis rearranges the same definition. Thus the ledger items are the entire input of the quantitative claims.

free parameters (9)
  • l (functionality coefficient) = 1
    Set to 1 in Table I to model a non-faulty discarded batch; this choice doubles the circularity penalty via mu_f,b = 1 + l in (11).
  • c_f,b,1 (criticality of beta1) = 0.1
    Chosen in Table I to represent a low-criticality material.
  • c_f,b,2 (criticality of beta2) = 0.95
    Chosen in Table I to represent a critical raw material.
  • t2,in,4 (time from extraction to disassembly facility) = 2,592,000 s (1 month)
    Arbitrary time constant in Table I.
  • Tr (reuse time) = 2,592,000 s (1 month)
    Arbitrary time constant in Table I.
  • Ti (incineration time) = 86,400 s (1 day)
    Arbitrary time constant in Table I.
  • Tt (transport time to incinerator) = 3,600 s (1 hour)
    Arbitrary time constant in Table I.
  • Delta (time conversion constant in (2)) = unspecified, must be fixed across calculations
    Defined in Definition 4 as an arbitrary constant to convert flow to mass; not used in this network because the continuous-flow term is zero.
  • Simulation step to real time conversion = 40 ms per step
    Taken from panda-gym [49]; used to convert trained controller step counts into Td (seconds).
assumptions (6)
  • domain assumption Compartmental dynamical thermodynamics framework (Haddad 2019)
    The entire TMN model (Definitions 1-4) is based on [18], treated as background.
  • domain assumption Mass conservation m0 = mu + mr
    Equation (7) imposes that extracted mass either is reused or is sent to incineration; no losses or other sinks.
  • ad hoc to paper Definition of finite-time sustainable mass (Definition 3)
    The metric rests on this definition of sustainability, which counts reservoir exit and sink entry as the same; this is the paper's own modeling choice.
  • ad hoc to paper Non-disassembled parts are considered functional (l=1)
    Equation (11) sets l=1 for all failed parts, maximizing the functionality penalty; a faulty batch would give l=0.
  • domain assumption Approximation tf ≈ t2,in,4 + Tr
    Section III-D states Td, Ti, Tt are much smaller than t2,in,4 and Tr; this approximation is used to derive (16)-(20).
  • domain assumption All parts weigh 1 kg and are made of beta1 or beta2
    Section III-C sets part masses to 1 kg for the simulation scenarios; the chassis adds an additional 3 kg in the hardest task.
invented entities (1)
  • Circular intelligence and robotics (CIRO)
    purpose: Name for an emerging research field combining circularity metrics with robotics and RL
    Introduced in the abstract and conclusions as a field label; no falsifiable predictions are attached.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CIRO7.2: A Material Network with Circularity of -7.2 and Reinforcement-Learning-Controlled Robotic Disassembler." pith.science (2026). https://pith.science/paper/TLMJX7LE

@misc{pith2026250611748,
  author       = {Pith},
  title        = {Pith review of: CIRO7.2: A Material Network with Circularity of -7.2 and Reinforcement-Learning-Controlled Robotic Disassembler},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TLMJX7LE}},
  note         = {Machine review of arXiv:2506.11748}
}
abstract

The competition over natural reserves of minerals is expected to increase in part because of the linear-economy paradigm based on take-make-dispose. Simultaneously, the linear economy considers end-of-use products as waste rather than as a resource, which results in large volumes of waste whose management remains an unsolved problem. Since a transition to a circular economy can mitigate these open issues, in this paper we begin by enhancing the notion of circularity based on compartmental dynamical thermodynamics, namely, $\lambda$, and then, we model a thermodynamical material network processing a batch of 2 solid materials of criticality coefficients of 0.1 and 0.95, with a robotic disassembler compartment controlled via reinforcement learning (RL), and processing 2-7 kg of materials. Subsequently, we focused on the design of the robotic disassembler compartment using state-of-the-art RL algorithms and assessing the algorithm performance with respect to $\lambda$ (Fig. 1). The highest circularity is -2.1 achieved in the case of disassembling 2 parts of 1 kg each, whereas it reduces to -7.2 in the case of disassembling 4 parts of 1 kg each contained inside a chassis of 3 kg. Finally, a sensitivity analysis highlighted that the impact on $\lambda$ of the performance of an RL controller has a positive correlation with the quantity and the criticality of the materials to be disassembled. This work also gives the principles of the emerging research fields indicated as circular intelligence and robotics (CIRO). Source code is publicly available.

Figures

Figures reproduced from arXiv: 2506.11748 by the authors.

Figure 1
Figure 1. (a) Disassembly task complexity vs. λ. (b) Final mean reward vs. λ. SAC-HER and TQC-HER are reinforcement-learning algorithms. An higher circularity λ ∈ (−∞, 0] indicates a more sustainable use of natural resources. the core of sustainability as defined by the United Nations Brundtland Commission back in 1987 [12]. To mitigate these fundamental issues existing both at na￾tional and international levels, the circular… view at source ↗
Figure 2
Figure 2. Compartmental diagraph of Ns (5). At the top, the description of each compartment; at the bottom, the masses and times are indicated with respect to the temporal axis. In green the disassembler compartment, which is considered in this paper. stage (c 2 2,2 ). The mass of non-renewable raw materials leaves the reservoir (c 1 1,1 ) at t = t1,out,4 ≡ 0. Consider that Ns processes 2 materials, namely, β1 and β2. Let m0 … view at source ↗
Figure 3
Figure 3. Visual simulation of disassembler compartment performance after training ( [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Further analysis of the robot behaviors after training (Figs. 4a and 4b) and during training (Figs. 4c and 4d). Specifically, Figs. 4a and 4b regard the [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 50 canonical work pages

  1. [17]

    Circular Economy Design through System Dynamics Modeling

    F. Zocco and M. Malvezzi, “Circular economy design through system dynamics modeling,”arXiv preprint arXiv:2411.13540, 2024

  2. [1]

    CRM Alliance, 2024, webpage: https://www.crmalliance.eu/ critical-raw-materials; last access: 28 April 2025

  3. [2]

    Critical raw materials,

    European Commission, “Critical raw materials,” 2023, webpage: https://single-market-economy.ec.europa.eu/sectors/raw-materials/ areas-specific-interest/critical-raw-materials en; last access: 28 April 2025

  4. [3]

    Notice of final determination on 2023 DOE critical materials list,

    Federal Register, “Notice of final determination on 2023 DOE critical materials list,” 2023, webpage: https: //www.federalregister.gov/documents/2023/08/04/2023-16611/ notice-of-final-determination-on-2023-doe-critical-materials-list; last access: 28 April 2025

  5. [4]

    Waste statistics,

    eurostat, “Waste statistics,” 2024, available at: https://ec.europa. eu/eurostat/statistics-explained/index.php?title=Waste statistics; last ac- cess: 28 April 2025

  6. [5]

    Critical minerals and great power com- petition: An overview,

    J. Zhou and A. M ˚anberger, “Critical minerals and great power com- petition: An overview,”STOCKHOLM INTERNATIONAL PEACE RE- SEARCH INSTITUTE, October 2024

  7. [6]

    A. Nygaard, “The geopolitical risk and strategic uncertainty of green growth after the Ukraine invasion: How the circular economy can decrease the market power of and resource dependency on critical minerals,”Circular Economy and Sustainability, vol. 3, no. 2, pp. 1099– 1126, 2023

  8. [7]

    Recycling and circular econ- omy—Towards a closed loop for metals in emerging clean technologies,

    C. Hagel ¨uken and D. Goldmann, “Recycling and circular econ- omy—Towards a closed loop for metals in emerging clean technologies,” Mineral Economics, vol. 35, no. 3, pp. 539–562, 2022

Show all 54 references
  1. [8]

    Towards circular economy implementation: a comprehensive review in context of manufacturing industry,

    M. Lieder and A. Rashid, “Towards circular economy implementation: a comprehensive review in context of manufacturing industry,”Journal of Cleaner Production, vol. 115, pp. 36–51, 2016

  2. [9]

    Demystifying series: Business and a circular economy,

    CE-Hub Team, “Demystifying series: Business and a circular economy,” 2022, available at: https://exeterce.org/knowledge-hub/ demystifying-business-and-a-circular-economy/# ftnref4; last access: 28 April 2025

  3. [10]

    Metal scarcity and sustain- ability, analyzing the necessity to reduce the extraction of scarce metals,

    M. Henckens, P. Driessen, and E. Worrell, “Metal scarcity and sustain- ability, analyzing the necessity to reduce the extraction of scarce metals,” Resources, Conservation and recycling, vol. 93, pp. 1–8, 2014

  4. [11]

    Scarce mineral resources: Extraction, consumption and limits of sustainability,

    T. Henckens, “Scarce mineral resources: Extraction, consumption and limits of sustainability,”Resources, Conservation and Recycling, vol. 169, p. 105511, 2021

  5. [12]

    Sustainability,

    United Nations, “Sustainability,” available at: https://www.un.org/en/ academic-impact/sustainability; last access: 28 April 2025

  6. [13]

    What is a circular economy?

    Ellen MacArthur Foundation, “What is a circular economy?” webpage: https://www.ellenmacarthurfoundation.org/topics/ circular-economy-introduction/overview; last access: 28 April 2025

  7. [14]

    Digitalisation for the transition to a resource efficient and circular economy,

    E. Bartekov ´a and P. B ¨orkey, “Digitalisation for the transition to a resource efficient and circular economy,”OECD Environment Working Papers, No. 192, OECD Publishing, Paris, 2022, https://doi.org/10.1787/ 6f6d18e7-en

  8. [15]

    Towards a thermodynamical deep-learning-vision-based flexible robotic cell for circular healthcare,

    F. Zocco, D. Sleath, and S. Rahimifard, “Towards a thermodynamical deep-learning-vision-based flexible robotic cell for circular healthcare,” Circular Economy and Sustainability, pp. 1–23, 2025

  9. [16]

    Introduction to an adaptive strategy for circular design,

    Ellen MacArthur Foundation, “Introduction to an adaptive strategy for circular design,” webpage: https://www.ellenmacarthurfoundation.org/ adaptive-strategy-for-circular-design/introduction; last access: 28 April 2025

  10. [18]

    W. M. Haddad,A Dynamical Systems Theory of Thermodynamics. Princeton University Press, 2019

  11. [19]

    Circular economy indicators for supply chains: A systematic literature review,

    T. Calzolari, A. Genovese, and A. Brint, “Circular economy indicators for supply chains: A systematic literature review,”Environmental and Sustainability Indicators, vol. 13, p. 100160, 2022

  12. [20]

    The future of circular economy metrics: Expert visions,

    M. Saidani, T. Shevchenko, Z. S. Esfandabadi, M. Ranjbari, J. A. Mesa, B. Yannou, and F. Cluzel, “The future of circular economy metrics: Expert visions,”Resources, Conservation and Recycling, vol. 205, p. 107565, 2024

  13. [21]

    The hidden concept and the beauty of multiple “R

    A. A. Zorpas, “The hidden concept and the beauty of multiple “R” in the framework of waste strategies development reflecting to circular economy principles,”Science of the Total Environment, vol. 952, p. 175508, 2024

  14. [22]

    Circular economy: Measuring innovation in the product chain,

    J. Potting, M. P. Hekkert, E. Worrell, and A. Hanemaaijer, “Circular economy: Measuring innovation in the product chain,”Planbureau voor de Leefomgeving, no. issue 2544, report, 2017

  15. [23]

    Developing a strategic methodology for circular economy roadmapping: A theoretical framework,

    H. Abu-Bakar and F. Charnley, “Developing a strategic methodology for circular economy roadmapping: A theoretical framework,”Sustainabil- ity, vol. 16, no. 15, p. 6682, 2024

  16. [24]

    Thermodynami- cal material networks for modeling, planning, and control of circular ma- terial flows,

    F. Zocco, P. Sopasakis, B. Smyth, and W. M. Haddad, “Thermodynami- cal material networks for modeling, planning, and control of circular ma- terial flows,”International Journal of Sustainable Engineering, vol. 16, no. 1, pp. 1–14, 2023

  17. [25]

    Circularity of thermodynamical material networks: Indicators, examples, and algorithms,

    F. Zocco, “Circularity of thermodynamical material networks: Indicators, examples, and algorithms,”arXiv preprint arXiv:2209.15051, 2024

  18. [26]

    Synchronized object detection for autonomous sorting, mapping, and quantification of JOURNAL OF LATEX CLASS FILES, VOL. 00, NO. 0, MONTH 20XX 8 materials in circular healthcare,

    F. Zocco, D. R. Lake, S. McLoone, and S. Rahimifard, “Synchronized object detection for autonomous sorting, mapping, and quantification of JOURNAL OF LATEX CLASS FILES, VOL. 00, NO. 0, MONTH 20XX 8 materials in circular healthcare,”arXiv preprint arXiv:2405.06821, 2024, to app...

  19. [27]

    Material flows and efficiency,

    J. M. Cullen and D. R. Cooper, “Material flows and efficiency,”Annual Review of Materials Research, vol. 52, no. 1, pp. 525–559, 2022

  20. [28]

    P. H. Brunner and H. Rechberger,Handbook of material flow analysis: For environmental, resource, and waste engineers. CRC press, 2016

  21. [29]

    A unification be- tween deep-learning vision, compartmental dynamical thermodynamics, and robotic manipulation for a circular economy,

    F. Zocco, W. M. Haddad, A. Corti, and M. Malvezzi, “A unification be- tween deep-learning vision, compartmental dynamical thermodynamics, and robotic manipulation for a circular economy,”IEEE Access, vol. 12, pp. 173 502–173 516, 2024

  22. [30]

    Robot for automatic waste sorting on construction sites,

    X. Chen, H. Huang, Y . Liu, J. Li, and M. Liu, “Robot for automatic waste sorting on construction sites,”Automation in Construction, vol. 141, p. 104387, 2022

  23. [31]

    Robotic waste sorting technology: Toward a vision- based categorization system for the industrial robotic separation of recyclable waste,

    M. Koskinopoulou, F. Raptopoulos, G. Papadopoulos, N. Mavrakis, and M. Maniadakis, “Robotic waste sorting technology: Toward a vision- based categorization system for the industrial robotic separation of recyclable waste,”IEEE Robotics & Automation Magazine, vol. 28, no. 2, pp...

  24. [32]

    Challenges for future robotic sorters of mixed industrial waste: a survey,

    T. Kiyokawa, J. Takamatsu, and S. Koyanaka, “Challenges for future robotic sorters of mixed industrial waste: a survey,”IEEE Transactions on Automation Science and Engineering, vol. 21, no. 1, pp. 1023–1040, 2022

  25. [33]

    Laili, Y

    Y . Laili, Y . Wang, Y . Fang, and D. T. Pham,Optimisation of robotic disassembly for remanufacturing. Springer, 2022

  26. [34]

    Enhancing disassembly practices for electric vehicle battery packs: a narrative comprehensive review,

    M. Beghi, F. Braghin, and L. Roveda, “Enhancing disassembly practices for electric vehicle battery packs: a narrative comprehensive review,” Designs, vol. 7, no. 5, p. 109, 2023

  27. [35]

    Robotic disassembly for increased recovery of strategically important materials from electrical vehicles,

    J. Li, M. Barwood, and S. Rahimifard, “Robotic disassembly for increased recovery of strategically important materials from electrical vehicles,”Robotics and Computer-Integrated Manufacturing, vol. 50, pp. 203–212, 2018. [Online]. Available: https://www.sciencedirect.com/ scie...

  28. [36]

    Robotic disassembly task training and skill transfer using reinforcement learning,

    M. Qu, Y . Wang, and D. T. Pham, “Robotic disassembly task training and skill transfer using reinforcement learning,”IEEE Transactions on Industrial Informatics, vol. 19, no. 11, pp. 10 934–10 943, 2023

  29. [37]

    Human- robot collaborative disassembly in industry 5.0: A systematic literature review and future research agenda,

    G. Yuan, X. Liu, X. Qiu, P. Zheng, D. T. Pham, and M. Su, “Human- robot collaborative disassembly in industry 5.0: A systematic literature review and future research agenda,”Journal of Manufacturing Systems, vol. 79, pp. 199–216, 2025

  30. [38]

    Transferring policy of deep reinforcement learning from simulation to reality for robotics,

    H. Ju, R. Juan, R. Gomez, K. Nakamura, and G. Li, “Transferring policy of deep reinforcement learning from simulation to reality for robotics,” Nature Machine Intelligence, vol. 4, no. 12, pp. 1077–1087, 2022

  31. [39]

    How to train your robot with deep reinforcement learning: lessons we have learned,

    J. Ibarz, J. Tan, C. Finn, M. Kalakrishnan, P. Pastor, and S. Levine, “How to train your robot with deep reinforcement learning: lessons we have learned,”The International Journal of Robotics Research, vol. 40, no. 4-5, pp. 698–721, 2021

  32. [40]

    A coach-based bayesian reinforcement learning method for snake robot control,

    Y . Jia and S. Ma, “A coach-based bayesian reinforcement learning method for snake robot control,”IEEE Robotics and Automation Letters, vol. 6, no. 2, pp. 2319–2326, 2021

  33. [41]

    Reinforcement learning models and algorithms for diabetes management,

    K.-L. A. Yau, Y .-W. Chong, X. Fan, C. Wu, Y . Saleem, and P.- C. Lim, “Reinforcement learning models and algorithms for diabetes management,”IEEE Access, vol. 11, pp. 28 391–28 415, 2023

  34. [42]

    Towards a multi-agent reinforcement learning approach for joint sensing and sharing in cognitive radio networks,

    K. Rapetswa and L. Cheng, “Towards a multi-agent reinforcement learning approach for joint sensing and sharing in cognitive radio networks,”Intelligent and Converged Networks, vol. 4, no. 1, pp. 50–75, 2023

  35. [43]

    Tc-driver: A trajectory conditioned reinforcement learning approach to zero-shot autonomous racing,

    E. Ghignone, N. Baumann, and M. Magno, “Tc-driver: A trajectory conditioned reinforcement learning approach to zero-shot autonomous racing,”Field Robotics, vol. 3, pp. 637–651, 2023

  36. [44]

    Reinforcement learning in robotic applications: a comprehensive survey,

    B. Singh, R. Kumar, and V . P. Singh, “Reinforcement learning in robotic applications: a comprehensive survey,”Artificial Intelligence Review, vol. 55, no. 2, pp. 945–990, 2022

  37. [45]

    J. A. Bondy and U. S. R. Murty,Graph Theory with Applications. Macmillan London, 1976, vol. 290

  38. [46]

    M. J. Moran, H. N. Shapiro, D. D. Boettner, and M. B. Bailey, Fundamentals of Engineering Thermodynamics. John Wiley & Sons, 2010

  39. [47]

    Thermodynamics: The unique universal science,

    W. M. Haddad, “Thermodynamics: The unique universal science,” Entropy, vol. 19, no. 11, p. 621, 2017

  40. [48]

    Circulate products and materials,

    Ellen MacArthur Foundation, “Circulate products and materials,” 2022, webpage: https://www.ellenmacarthurfoundation.org/ circulate-products-and-materials; last access: 17 April 2025

  41. [49]

    panda-gym: Open-source goal-conditioned environments for robotic learning,

    Q. Gallou ´edec, N. Cazin, E. Dellandr ´ea, and L. Chen, “panda-gym: Open-source goal-conditioned environments for robotic learning,” in 4th Robot Learning Workshop: Self-Supervised and Lifelong Learning@ NeurIPS 2021, 2021

  42. [50]

    Stable-baselines3: Reliable reinforcement learning implementa- tions,

    A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dor- mann, “Stable-baselines3: Reliable reinforcement learning implementa- tions,”Journal of Machine Learning Research, vol. 22, no. 268, pp. 1–8, 2021

  43. [51]

    Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

    T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,” inInternational Conference on Machine Learning. Pmlr, 2018, pp. 1861–1870

  44. [52]

    Controlling overestimation bias with truncated mixture of continuous distributional quantile critics,

    A. Kuznetsov, P. Shvechikov, A. Grishin, and D. Vetrov, “Controlling overestimation bias with truncated mixture of continuous distributional quantile critics,” inInternational Conference on Machine Learning. PMLR, 2020, pp. 5556–5566

  45. [53]

    Addressing function approxi- mation error in actor-critic methods,

    S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approxi- mation error in actor-critic methods,” inInternational Conference on Machine Learning. PMLR, 2018, pp. 1587–1596

  46. [54]

    Hindsight experience replay,

    M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba, “Hindsight experience replay,”Advances in Neural Information Processing Systems, vol. 30, 2017

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.