REVIEW 5 major objections 4 minor 64 references
A Study of the Efficacy of Generative Flow Networks for Robotics and Machine Fault-Adaptation
T0 review · 5 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper claims that continuous generative flow networks (CFlowNets) can adapt a simulated two-joint robot arm to four injected faults faster and, in most cases, with higher asymptotic reward than DDPG, TD3, SAC, and PPO, at the cost of…
desk verdict A genuinely new application of CFlowNets to fault adaptation, but the headline comparison is confounded by reward asymmetry and a pre-trained retrieval network; the conclusion overreaches. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is flow matching over continuous trajectories: CFlowNets parameterizes an edge flow function $F_\theta(s,a)$ that approximates how much probability mass flows through each state-action edge, then trains it so that the inflow to a state equals the outflow plus the terminal reward, with a retrieval network $G_\phi$ approximating parent states. At decision time it uniformly samples $M$ candidate actions, scores them with $F_\theta$, and samples an action with probability proportional to that score. This keeps a distribution over many high-reward paths alive instead of collapsing onto a single policy, which is the property the paper credits for fast adaptation; the cost is that every step scores $M$ actions and matches flows over $K$ sampled actions, driving the measured memory and compute overhead.
What would settle it
Run all five agents in the same four fault environments with one shared reward scheme and equal hyperparameter budgets, and train CFlowNets' retrieval network from scratch during fault adaptation; if CFlowNets no longer reaches asymptotic reward in roughly 100,000 to 200,000 timesteps or no longer matches PPO's final reward, the paper's central claim fails.
Extended reading notes
Core claim
On its own terms, the paper discovers that CFlowNets can be trained on a normal reaching task and then re-adapt to hardware faults with far fewer environmental interactions than DDPG, TD3, SAC, or PPO. In four fault environments created by editing the simulator configuration, CFlowNets reaches its asymptotic reward in as few as 0.1 million timesteps and records the highest asymptotic reward in three of the four cases. The paper's conclusion states that CFlowNets outperforms state-of-the-art RL algorithms in the Reacher-v2 robotic environment, while the Discussion acknowledges that CFlowNets operated with a pre-trained retrieval network and a sparse reward, whereas the RL agents received dense per-step rewards and less tuning.
Load-bearing premise
The conclusion rests on the assumption that CFlowNets and the RL algorithms were compared on equal footing; the paper's own Discussion concedes that CFlowNets had a pre-trained helper network and saw a reward only at the end of each attempt, while the RL agents got a reward at every step and less tuning, so the measured speed advantage may not be solely about the algorithms.
Editorial extensions
If this is right
- CFlowNets reaches near-asymptotic reward within roughly 0.1 to 0.2 million timesteps for three fault environments, while PPO needs 3.5 to 4.6 million and DDPG needs 5.5 to 6.1 million timesteps.
- CFlowNets ends with higher asymptotic reward than all RL baselines in the reduced-range, actuator-damage, and structural-damage environments, and lands close to PPO in the increased-damping environment.
- Retaining both the pre-fault model and replay buffer gives a jumpstart in three fault environments but hurts performance in the reduced-range-of-motion environment.
- CFlowNets retains 68 to 95 percent of its normal reward in three fault environments but drops to about 21 percent retention under actuator damage, showing that the advantage is fault-specific.
- The fast adaptation comes with a compute cost: CFlowNets averaged 17.91 GB of GPU memory and roughly 5 hours 39 minutes per million timesteps, while the RL baselines used under one-third of that memory.
- CFlowNets is positioned as a method for exploration-biased tasks where many good solutions exist, while the paper concedes that traditional RL may be more suitable when the goal is strictly maximizing cumulative reward.
Reading between the lines
- If the reward asymmetry were removed, the comparison would be the decisive test: giving CFlowNets the same dense per-step reward, or reducing the RL agents to terminal-only rewards, could shrink or enlarge the measured speed gap, and the paper's own Discussion implies neither condition was tested.
- The mechanism suggests CFlowNets would be most valuable in fault scenarios with many viable compensatory strategies, where sampling a distribution over high-reward paths beats committing to a single policy; a coupled-fault benchmark combining actuator damage with increased damping would test this directly.
- The 17.91 GB memory footprint implies that making CFlowNets practical for embedded robots would require distillation, pruning, or an approximate flow parameterization; otherwise the method is confined to server-side training with deployment to cheaper policies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies whether continuous generative flow networks (CFlowNets) can provide fast fault adaptation in a simulated robotic task. The authors modify the MuJoCo Reacher-v2 environment to create four fault conditions (reduced range of motion, increased damping, actuator damage, structural damage), compare CFlowNets against DDPG, TD3, SAC, and PPO over 10 million timesteps, and report asymptotic reward (Table 2), adaptation speed (Figure 7, Table 3), execution time (Figure 8), GPU memory (Figure 9), and a transfer-learning analysis for CFlowNets (Figure 10). The main claims are that CFlowNets adapts with the fewest samples, achieves high asymptotic performance, and 'outperforms state-of-the-art RL algorithms' in Reacher-v2 (Section 7). The paper is transparent about several limitations, but those limitations directly affect the validity of the headline comparison.
Significance. If the central comparison were fair, this would be a useful first demonstration of GFlowNets for robot fault adaptation, with a reproducible public codebase, ten independent runs per condition, and honest reporting of compute costs. The paper also makes a sensible distinction between adaptation speed and asymptotic performance, and it reports resource usage rather than only rewards. However, the main significance hinges on the comparison being apples-to-apples, and the paper itself discloses asymmetries in reward structure, pretrained components, and hyperparameter tuning that undermine that assumption. The contribution is therefore valuable as a preliminary study, but the headline conclusion substantially overstates what the evidence supports.
major comments (5)
- [Section 6 (Discussion)] The reward structure for CFlowNets and the RL baselines was not the same. The Discussion states that CFlowNets 'operated in a sparse reward-structured environment whereas the RL algorithms received intermediate rewards,' with CFlowNets receiving a reward only at the terminal state. This is not a minor tuning detail: the reward function is part of the MDP, so the algorithms were optimizing different objectives. Consequently, the asymptotic reward values in Table 2 and the adaptation-speed timesteps in Figure 7 and Table 3 compare different tasks, and the Section 7 conclusion that CFlowNets 'outperforms state-of-the-art RL algorithms' is not established. The authors suggest the sparse reward was a disadvantage that roughly balances CFlowNets' pretrained retrieval network, but no experiment supports that balancing claim. A fair comparison with identical reward structure, or a reframing of the results as comparing two different problem formulations, is required.
- [Section 5.3 and Section 6] CFlowNets' reported execution time excludes the pretraining of the retrieval network, which is a substantial component of the method. Section 5.3 notes that the reported execution time only accounts for training the flow network with an already pretrained retrieval network, and Section 6 acknowledges that this pretrained component may explain why CFlowNets 'was able to quickly gain convergence compared to other RL algorithms.' No RL baseline received an equivalent pretrained component. Since this asymmetry points in the opposite direction from the sparse-reward asymmetry, the net effect on the comparison is unknown. The authors should either include the retrieval pretraining time in the resource comparison, train the retrieval network jointly, or provide an ablation that quantifies the contribution of the pretrained retrieval network.
- [Section 4.5 and Section 6] The hyperparameter tuning effort was asymmetric. CFlowNets received a dedicated hyperparameter search described in Section 4.5, while the Discussion states that 'an intensive hyperparameter search was not conducted for the implementation of the RL algorithms' and that published hyperparameters with selective exploration tuning were used. This is particularly relevant for DDPG, which the paper itself notes is highly sensitive to hyperparameters. The 'outperforms' claim therefore conflates algorithmic capability with tuning effort. A fair comparison should either run the RL baselines with a comparable tuning budget or explicitly present the results as 'CFlowNets with tuned hyperparameters versus RL baselines with default hyperparameters.'
- [Section 7 and Table 2] The conclusion that CFlowNets 'achieved a high asymptotic performance, surpassing the state-of-the-art RL algorithms' is contradicted by the paper's own Table 2. In the Increased Damping environment, PPO reaches -4.7 versus CFlowNets' -4.9, and in the Actuator Damage environment, PPO reaches -6.5 versus CFlowNets' -6.8. The text in Section 5.1 also states that PPO 'outperforming every algorithm in terms of higher asymptotic performance for the Fault 3 environment.' Thus the asymptotic-performance claim in the conclusion is too strong even under the authors' own numbers. Moreover, no statistical significance tests accompany these small differences, so claims of superiority in asymptotic reward are not supported. The authors should either soften the conclusion to 'comparable or better asymptotic performance in most environments' or provide a significance analysis.
- [Section 5.5 and Figure 10] The transfer experiment is difficult to interpret because the text and the figure legend disagree. Section 5.5 says the comparison is between retaining both model parameters and replay buffer versus retaining 'only the model parameters,' but Figure 10 labels the curves as 'No Prior Learning' and 'Retained Model and Storage.' It is unclear whether 'No Prior Learning' means training from scratch, retaining only the model, or something else. In addition, the 'performance retention' percentages in Section 6 (e.g., 68.43% to 94.74%) reference the asymptotic performance in the normal environment, but that normal-environment baseline is not reported in a table or figure alongside the fault results. These omissions make the fourth contribution hard to verify.
minor comments (4)
- [Throughout] There are numerous typos and grammatical errors, including 'Conlcusion' in the Section 7 heading, 'Schamatics' in the Figure 3 caption, 'is comes at a computational cost' in Section 5.3, 'unforseen' in the abstract, and 'Rettained' in the Figure 10 caption. A careful proofreading pass is needed.
- [Section 4.1.1] The environment name is spelled inconsistently as 'Mujoco' and 'MuJoCo'; the latter is the correct spelling used by the simulator itself.
- [Section 2.2] The notation for the self-conditional flow function F(s|s') is introduced but never used afterward; the paper should either use it in the loss derivation or omit it.
- [Table 3] Table 3 is difficult to read in the submitted format because the header and entries are packed together without clear column separation; the authors should reformat the table with explicit column headers.
Circularity Check
No significant circularity: the paper is an empirical benchmark study; its comparisons are vulnerable to fairness concerns but not to derivational self-reference.
full rationale
The paper does not derive its headline results from its inputs by construction. CFlowNets is trained with the published flow-matching loss (Eq. 9) against the environment reward, and adaptation speed and asymptotic reward are measured from rollouts; no fitted parameter is renamed as a prediction, and no equation is defined in terms of the quantity it claims to establish. The fault severities are selected by an attribute search, which affects comparability, but that is experimental design rather than circular reasoning. The acknowledged asymmetries (sparse terminal reward for CFlowNets versus dense intermediate rewards for RL, and the pre-trained retrieval network) are validity threats to the 'outperforms' claim, not circular reductions. The only self-citation used to justify omitting certain transfer configurations ([53]) is not load-bearing for the central claim and does not replace an independent derivation. The evaluation is self-contained against external baselines, so no circularity score above zero is warranted.
Assumptions & free parameters
free parameters (3)
- Fault severity values =
damping=5, actuator power=100, joint1 range=[-1,1], link1 bend=45 degrees
- Asymptotic performance variance threshold =
not specified
- CFlowNets hyperparameters =
lr=0.003, batch=256, K=100, action buffer=10,000, replay buffer=100,000
assumptions (3)
- domain assumption CFlowNets training converges to a distribution over trajectories proportional to reward.
- domain assumption The four XML attribute changes faithfully emulate real-world mechanical faults.
- domain assumption Sparse (terminal-only) rewards for CFlowNets and dense rewards for RL baselines produce comparable training signals.
Cite this review
Pith. "Pith review of A Study of the Efficacy of Generative Flow Networks for Robotics and Machine Fault-Adaptation." pith.science (2026). https://pith.science/paper/GWIQ7AJL
@misc{pith2026250103405,
author = {Pith},
title = {Pith review of: A Study of the Efficacy of Generative Flow Networks for Robotics and Machine Fault-Adaptation},
year = {2026},
howpublished = {\url{https://pith.science/paper/GWIQ7AJL}},
note = {Machine review of arXiv:2501.03405}
}
read the original abstract
Advancements in robotics have opened possibilities to automate tasks in various fields such as manufacturing, emergency response and healthcare. However, a significant challenge that prevents robots from operating in real-world environments effectively is out-of-distribution (OOD) situations, wherein robots encounter unforseen situations. One major OOD situations is when robots encounter faults, making fault adaptation essential for real-world operation for robots. Current state-of-the-art reinforcement learning algorithms show promising results but suffer from sample inefficiency, leading to low adaptation speed due to their limited ability to generalize to OOD situations. Our research is a step towards adding hardware fault tolerance and fast fault adaptability to machines. In this research, our primary focus is to investigate the efficacy of generative flow networks in robotic environments, particularly in the domain of machine fault adaptation. We simulated a robotic environment called Reacher in our experiments. We modify this environment to introduce four distinct fault environments that replicate real-world machines/robot malfunctions. The empirical evaluation of this research indicates that continuous generative flow networks (CFlowNets) indeed have the capability to add adaptive behaviors in machines under adversarial conditions. Furthermore, the comparative analysis of CFlowNets with reinforcement learning algorithms also provides some key insights into the performance in terms of adaptation speed and sample efficiency. Additionally, a separate study investigates the implications of transferring knowledge from pre-fault task to post-fault environments. Our experiments confirm that CFlowNets has the potential to be deployed in a real-world machine and it can demonstrate adaptability in case of malfunctions to maintain functionality.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
Z. Sun, H. Yang, Y. Ma, X. Wang, Y. Mo, H. Li, Z. Jiang, Bit-dmr: A humanoid dual-arm mobile robot for complex rescue operations, IEEE Robotics and Automation Letters 7 (2) (2022) 802–809. doi:10.1109/ LRA.2021.3131379
arXiv 2022
-
[3]
O. SeungSub, H. Jehun, J. Hyunjung, L. Soyeon, S. Jinho, A study on the disaster response scenarios using robot technology, in: 2017 14th In- ternational Conference on Ubiquitous Robots and Ambient Intelligence (URAI), 2017, pp. 520–523. doi:10.1109/URAI.2017.7992658
-
[4]
J. Burgner-Kahrs, D. C. Rucker, H. Choset, Continuum robots for medi- cal applications: A survey, IEEE Transactions on Robotics 31 (6) (2015) 1261–1280. doi:10.1109/TRO.2015.2489500
arXiv 2015
-
[5]
P. Papadakis, Terrain traversability analysis methods for unmanned ground vehicles: A survey, Engineering Applications of Artificial In- telligence 26 (4) (2013) 1373–1385
work page 2013
-
[6]
L. Liu, P. Li, Plant intelligence-based pillo underwater target detection algorithm, Engineering Applications of Artificial Intelligence 126 (2023) 106818
work page 2023
-
[7]
A. S. Chen, G. Chada, L. Smith, A. Sharma, Z. Fu, S. Levine, C. Finn, Adapt on-the-go: Behavior modulation for single-life robot deployment, arXiv preprint arXiv:2311.01059 (2023)
arXiv 2023
-
[8]
M. Fernandes, J. M. Corchado, G. Marreiros, Machine learning tech- niques applied to mechanical fault diagnosis and fault prognosis in the context of real industrial manufacturing use-cases: a systematic litera- ture review, Applied Intelligence 52 (12) (2022) 14246–14280
work page 2022
Show all 64 references
-
[9]
R. Teti, K. Jemielniak, G. O’Donnell, D. Dornfeld, Advanced monitoring of machining operations, CIRP annals 59 (2) (2010) 717–739. 46
2010
-
[10]
Riazi, O
M. Riazi, O. Zaiane, T. Takeuchi, A. Maltais, J. G¨ unther, M. Lipsett, Detecting the onset of machine failure using anomaly detection meth- ods, in: Big Data Analytics and Knowledge Discovery: 21st Interna- tional Conference, DaWaK 2019, Linz, Austria, August 26–29, 2019, Pro...
2019
-
[11]
Castellano-Quero, M
M. Castellano-Quero, M. Castillo-L´ opez, J.-A. Fern´ andez-Madrigal, V. Ar´ evalo-Espejo, H. Voos, A. Garc ´ ıa-Cerezo, A multidimensional bayesian architecture for real-time anomaly detection and recovery in mobile robot sensory systems, Engineering Applications of Artificia...
2023
-
[12]
Chatti, R
N. Chatti, R. Guyonneau, L. Hardouin, S. Verron, S. Lagrange, Model- based approach for fault diagnosis using set-membership formulation, Engineering Applications of Artificial Intelligence 55 (2016) 307–319
2016
-
[13]
Guiochet, M
J. Guiochet, M. Machin, H. Waeselynck, Safety-critical advanced robots: A survey, Robotics and Autonomous Systems 94 (2017) 43–52
2017
-
[14]
Afzaal, J.-A
U. Afzaal, J.-A. Lee, Low-cost hardware redundancy for fault-mitigation in power-constrained iot systems, in: 2020 International Conference on Information and Communication Technology Convergence (ICTC), 2020, pp. 60–62. doi:10.1109/ICTC49870.2020.9289420
2020
-
[15]
Frank, A
F. Frank, A. Paraschos, P. van der Smagt, B. Cseke, Constrained prob- abilistic movement primitives for robot trajectory adaptation, IEEE Transactions on Robotics 38 (4) (2022) 2276–2294. doi:10.1109/TRO. 2021.3127108
2022
-
[16]
Z. Luo, E. Xiao, P. Lu, Ft-net: Learning failure recovery and fault- tolerant locomotion for quadruped robots, IEEE Robotics and Automa- tion Letters 8 (12) (2023) 8414–8421. doi:10.1109/LRA.2023.3329766
2023
-
[17]
The Engineer, Forces of nature: Biomimicry in robotics, The Engineer- Accessed: 2024-01-15 (2015)
2015
-
[18]
Cully, J
A. Cully, J. Clune, D. Tarapore, J.-B. Mouret, Robots that can adapt like animals, Nature 521 (7553) (2015) 503–507
2015
-
[19]
X. Song, Y. Yang, K. Choromanski, K. Caluwaerts, W. Gao, C. Finn, J. Tan, Rapidly adaptable legged robots via evolutionary meta-learning, 47 in: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2020, pp. 3769–3776
2020
-
[20]
R. S. Sutton, A. G. Barto, Reinforcement learning: An introduction, MIT press, 2018
2018
-
[21]
Arulkumaran, M
K. Arulkumaran, M. P. Deisenroth, M. Brundage, A. A. Bharath, A brief survey of deep reinforcement learning, arXiv preprint arXiv:1708.05866 (2017)
2017 arXiv
-
[22]
S. Gu, E. Holly, T. P. Lillicrap, S. Levine, Deep reinforcement learning for robotic manipulation, arXiv preprint arXiv:1610.00633 1 (2016) 1
2016 arXiv
-
[23]
Raziei, M
Z. Raziei, M. Moghaddam, Adaptable automation with modular deep reinforcement learning and policy transfer, Engineering Applications of Artificial Intelligence 103 (2021) 104296
2021
-
[24]
Bengio, S
Y. Bengio, S. Lahlou, T. Deleu, E. J. Hu, M. Tiwari, E. Bengio, Gflownet foundations, Journal of Machine Learning Research 24 (210) (2023) 1– 55
2023
-
[25]
Y. Li, S. Luo, H. Wang, J. Hao, Cflownets: Continuous control with generative flow networks, International Conference on Learning Repre- sentations (2023)
2023
-
[26]
D. W. Zhang, C. Rainone, M. Peschl, R. Bondesan, Robust scheduling with gflownets, arXiv preprint arXiv:2302.05446 (2023)
2023 arXiv
-
[27]
M. W. Shen, E. Bengio, E. Hajiramezanali, A. Loukas, K. Cho, T. Bian- calani, Towards understanding and improving gflownet training, in: In- ternational Conference on Machine Learning, PMLR, 2023, pp. 30956– 30975
2023
-
[28]
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Sil- ver, D. Wierstra, Continuous control with deep reinforcement learning, 4th International Conference on Learning Representations, ICLR, San Juan, Puerto Rico, May 2-4, Conference Track Proceedings (2016)
2016
-
[29]
Fujimoto, H
S. Fujimoto, H. Hoof, D. Meger, Addressing function approximation error in actor-critic methods, in: International conference on machine learning, PMLR, 2018, pp. 1587–1596. 48
2018
-
[30]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal policy optimization algorithms, CoRR (2017)
2017
-
[31]
Haarnoja, A
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Ku- mar, H. Zhu, A. Gupta, P. Abbeel, et al., Soft actor-critic algorithms and applications, CoRR (2018). URL http://arxiv.org/abs/1812.05905
2018 arXiv
-
[32]
M. T. Rosenstein, Z. Marx, L. P. Kaelbling, T. G. Dietterich, To transfer or not to transfer, in: Proceedings of the NIPS 2005 workshop on transfer learning, NIPS, Whistler, British Columbia, 2005, pp. 1–4
2005
-
[33]
C. Finn, K. Xu, S. Levine, Probabilistic model-agnostic meta-learning, Advances in neural information processing systems 31 (2018)
2018
-
[34]
R. S. Sutton, D. McAllester, S. Singh, Y. Mansour, Policy gradient methods for reinforcement learning with function approximation, Ad- vances in neural information processing systems 12 (1999)
1999
-
[35]
Malkin, M
N. Malkin, M. Jain, E. Bengio, C. Sun, Y. Bengio, Trajectory balance: Improved credit assignment in gflownets, CoRR abs/2201.13259 (2022). arXiv:2201.13259. URL https://arxiv.org/abs/2201.13259
2022 arXiv
-
[36]
Bengio, K
Y. Bengio, K. Malkin, M. Jain, The gflownet tutorial (2022). URL https://tinyurl.com/gflownet-tutorial
2022
-
[37]
Chatzilygeroudis, V
K. Chatzilygeroudis, V. Vassiliades, J.-B. Mouret, Reset-free trial-and- error learning for robot damage recovery, Robotics and Autonomous Systems 100 (2018) 236–250
2018
-
[38]
Ahmed, M
I. Ahmed, M. Quinones-Grueiro, G. Biswas, Complementary meta- reinforcement learning for fault-adaptive control, arXiv preprint arXiv:2009.12634 (2020)
2020 arXiv
-
[39]
M. Luo, A. Balakrishna, B. Thananjeyan, S. Nair, J. Ibarz, J. Tan, C. Finn, I. Stoica, K. Goldberg, Mesa: Offline meta-rl for safe adaptation and fault tolerance, arXiv preprint arXiv:2112.03575 (2021)
2021 arXiv
-
[40]
Nagabandi, I
A. Nagabandi, I. Clavera, S. Liu, R. S. Fearing, P. Abbeel, S. Levine, C. Finn, Learning to adapt in dynamic, real-world environments through meta-reinforcement learning (2019). arXiv:1803.11347. 49
2019 arXiv
-
[41]
C. Finn, P. Abbeel, S. Levine, Model-agnostic meta-learning for fast adaptation of deep networks, in: International conference on machine learning, PMLR, 2017, pp. 1126–1135
2017
-
[42]
Fournier, O
P. Fournier, O. Sigaud, M. Chetouani, P.-Y. Oudeyer, Accuracy-based curriculum learning in deep reinforcement learning, arXiv preprint arXiv:1806.09614 (2018)
2018 arXiv
-
[43]
Parisotto, S
E. Parisotto, S. Ghosh, S. B. Yalamanchi, V. Chinnaobireddy, Y. Wu, R. Salakhutdinov, Concurrent meta reinforcement learning, arXiv preprint arXiv:1903.02710 (2019)
2019 arXiv
-
[44]
Brockman, V
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, W. Zaremba, Openai gym (2016). arXiv:arXiv:1606.01540
2016 arXiv
-
[45]
M. C. Tips, Environmental effects on motion components in robots, Motion Control Tips (2023). URL https://www.motioncontroltips.com
2023
-
[46]
Robots, Working of robotic arm: How does a robotic arm work?, accessed: 2023-04-15 (2022)
U. Robots, Working of robotic arm: How does a robotic arm work?, accessed: 2023-04-15 (2022). URL https://www.universal-robots.com
2022
-
[47]
Robots, About universal robots, accessed: 2023-04-15 (2022)
U. Robots, About universal robots, accessed: 2023-04-15 (2022). URL https://www.universal-robots.com
2022
-
[48]
Limited, Friction model of industrial robot joint with temperature correction by example of kuka kr10, Hindawi Publishing Corporation (2023)
H. Limited, Friction model of industrial robot joint with temperature correction by example of kuka kr10, Hindawi Publishing Corporation (2023). URL https://www.hindawi.com
2023
-
[49]
A. C. Bittencourt, Static friction in a robot joint—modeling and iden- tification of load dependencies, ASME Digital Collection (2023). URL https://asmedigitalcollection.asme.org
2023
-
[50]
Ferretti, G
G. Ferretti, G. Magnani, P. Rocco, Modelling the temperature in joint friction of industrial manipulators, Robotica, Cambridge University Press (2023). URL https://www.cambridge.org 50
2023
-
[51]
I. A. Co., 12 causes of servo motor failure, Industrial Automation InsightsAccessed: 2024-04-15 (2023). URL https://www.industrialautomationco.com/ servo-motor-failure
2023
-
[52]
America, 5 servo motor failure causes & how to prevent them, KEB America InsightsAccessed: 2024-04-15 (2023)
K. America, 5 servo motor failure causes & how to prevent them, KEB America InsightsAccessed: 2024-04-15 (2023). URL https://www.kebamerica.com/servo-motor-failures
2023
-
[53]
Schoepp, M
S. Schoepp, M. Taghian, S. Miwa, Y. Mitsuka, S. Golestan, O. Za ¨ ıane, Enhancing hardware fault tolerance in machines with reinforcement learning policy gradient algorithms, arXiv preprint arXiv:2407.15283 (2024)
2024 arXiv
-
[54]
Yarats, I
D. Yarats, I. Kostrikov, Soft actor-critic (sac) implementation in py- torch, https://github.com/denisyarats/pytorch_sac (2020)
2020
-
[55]
Raffin, A
A. Raffin, A. Hill, M. Ernestus, A. Gleave, A. Kanervisto, Stable baselines3: Reliable reinforcement learning implementations (2021). URL https://www.jmlr.org/papers/volume22/20-1364/20-1364. pdf
2021
-
[56]
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, K. Kavukcuoglu, Asynchronous methods for deep reinforce- ment learning (2016). arXiv:1602.01783
2016 arXiv
-
[57]
Niroui, K
F. Niroui, K. Zhang, Z. Kashino, G. Nejat, Deep reinforcement learn- ing robot for search and rescue applications: Exploration in unknown cluttered environments, IEEE Robotics and Automation Letters 4 (2) (2019) 610–617
2019
-
[58]
M. E. Taylor, P. Stone, Transfer Learning for Reinforcement Learning Domains: A Survey, Journal of Machine Learning Research 10 (1) (2009) 1633–1685
2009
-
[59]
Marculescu, D
D. Marculescu, D. Stamoulis, E. Cai, Hardware-aware machine learning: Modeling and optimization, in: 2018 IEEE/ACM International Confer- ence on Computer-Aided Design (ICCAD), IEEE, 2018, pp. 1–8. 51
2018
-
[60]
Paleyes, R.-G
A. Paleyes, R.-G. Urma, N. D. Lawrence, Challenges in deploying ma- chine learning: a survey of case studies, ACM computing surveys 55 (6) (2022) 1–29
2022
-
[61]
Queeney, I
J. Queeney, I. C. Paschalidis, C. G. Cassandras, Generalized Proximal Policy Optimization with Sample Reuse, Advances in Neural Informa- tion Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS, December 6-14, virtual (2021)
2021
-
[62]
Henderson, R
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, D. Meger, Deep reinforcement learning that matters, Proceedings of the Thirty- Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the...
2018
-
[63]
Matheson, R
E. Matheson, R. Minto, E. G. Zampieri, M. Faccio, G. Rosati, Human– robot collaboration in manufacturing applications: A review, Robotics 8 (4) (2019) 100
2019
-
[64]
M. C. Capolei, E. Angelidis, E. Falotico, H. H. Lund, S. Tolu, A biomimetic control method increases the adaptability of a humanoid robot acting in a dynamic environment, Frontiers in neurorobotics 13 (2019) 70. 52
2019
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.