:
twitter line
研究生: 尤思涵
研究生(外文): Szu-HanYu
論文名稱: 整合模糊控制改良深度決定性決策梯度網路及粒子群演算法實現雙手服務型機器人之手臂動作規劃與控制
論文名稱(外文): Integration of FLC improved DDPG and PSO for Motion Planning and Control of Dual-Arm Home Service Robot
指導教授: 李祖聖
指導教授(外文): Tzuu-Hseng S. Li
學位類別: 碩士
校院名稱: 國立成功大學
系所名稱: 電機工程學系
學門: 工程學門
學類: 電資工程學類
論文種類: 學術論文
論文出版年: 2020
畢業學年度: 108
語文別: 英文
論文頁數: 127
中文關鍵詞: 深度決定性決策梯度網路 模糊邏輯控制 運動軌跡規劃 粒子演算法
外文關鍵詞: Deep Deterministic Policy Gradients Fuzzy Logic Control Motion Planning Particle Swarm Optimization
相關次數:
  • 被引用 被引用:0
  • 點閱 點閱:204
  • 評分 評分:
  • 下載 下載:0
  • 收藏至我的研究室書目清單 書目收藏:1
機器人雙手手臂運動軌跡規劃,係經由規劃雙手末端點的移動軌跡及相對應的各軸馬達角度,以控制雙手在空間中之移動。為了人類與機構的安全性,讓雙手在移動過程中避障,是一個重要的議題,由於七軸手臂的逆運動學有無限多組解,如何挑選合適的解,也是困難的問題。為了解決上述議題,本論文提出一個整合粒子群演算法(PSO)與模糊邏輯控制改良之深度決定性決策梯度網路(Fuzzy-DDPG)的雙手手臂運動軌跡規劃系統。運用順向運動學得到雙手的狀態特徵,於模擬環境進行互動,經由避障獎懲函數之設計,使手臂於深度決定性決策梯度網路所產生的動作與評價結果中進行學習,學習最適合當時情況的馬達轉角序列,並在空間中形成軌跡,此系統能同時考慮移動軌跡及轉角,避免軌跡中點與點之間的不確定性。此外,運用模糊邏輯控制對網路的更新部分進行改良,除了軌跡的第一點,在其餘的軌跡點上,皆會藉由模糊邏輯及考慮該軌跡點在軌跡中的序位,去決定是否要進行探索,進而使網路中資料庫的資料更多樣化,讓網路可以收斂更快跳出區域最佳解,規劃更好的移動路徑。為了改善深度決定性決策梯度網路較不適用於隨機環境的缺點,本論文運用粒子群演算法將經由網路所規劃出的軌跡進行優化,提高軌跡的避障能力及使軌跡更加平滑。本論文除了在模擬環境中驗證所提方法外,也執行數個實際實驗場景,包含單手與雙手的移動以及雙手間的互動,展示了本論文所提方法能有效地控制雙手手臂之運動。
Motion planning for the dual arms with high degree of freedom is a critical research topic for home service robots. For the safety of human being, the problem of collision avoidance during these two arms manipulate in the environment is a very important issue. The solution of Inverse Kinematics for the 7-DoF is infinite; therefore, how to select the proper solution is also a real challenge. In order to cope with the problems, this thesis proposes a Fuzzy-Deep Deterministic Policy Gradients (Fuzzy-DDPG) with Particle Swarm Optimization (PSO) improvement method applied to the dual 7-DoF manipulators. Through the information from the forward kinematics and the designed reward according to the potential collision, the manipulators are able to learn the suitable trajectories and the corresponding motor rotation angles from the generated actions and performance judgment by DDPG. Furthermore, this thesis integrates DDPG with the concept of Fuzzy Logic Control (FLC) to improve the performance by updating the network weights. The Fuzzy-DDPG method successfully enlarges the exploration area of trajectory points, decreases the probability of local minimum, and increases the speed of network convergence. In addition, for the sake of the difference between simulation and real environment, this thesis also adopts PSO algorithm to slightly modify the generated trajectories, which guarantees the collision avoidance and smooth movement. By verifying the validation in the simulations and real-time experiments, we set up several scenarios and apply the Fuzzy-DDPG with PSO trajectory optimization scheme to the dual-arm home service robot in the environment of obstacles. The results successfully demonstrate that the method can control the dual arms to arrive the destination in a smooth trajectory without any collision.
Abstract Ⅰ
Acknowledgement ⅠⅠⅠ
Contents ⅠⅤ
List of Figures Ⅶ
List of Tables Ⅻ
Chapter 1 Introduction 1
1.1 Motivation 1
1.2 Related Works 2
1.3 System Overview 4
1.4 Thesis Organization 5
Chapter 2 Robot Hardware and Simulated Environment 6
2.1 Introduction 6
2.2 Dual-Arm Hardware and Kinematics 8
2.3 Simulation Environment 12
Chapter 3 Motion Planning System Based on DDPG 15
3.1 Introduction 15
3.2 General Structure of DDPG 17
3.2.1 Data Collection from Environmental Interaction 18
3.2.2 Weight Update Method 20
3.3 DDPG Design Applied to Dual-Arm Manipulators 25
3.3.1 State Design for Dual Arms 25
3.3.2 Action Design for Motion 28
3.3.3 Network Design for Training 29
3.3.4 Reward Design for Learning 32
3.4 DDPG Experiments in Simulated Environment 36
3.4.1 Training 37
3.4.2 Testing Results 53
Chapter 4 Fuzzy-DDPG Method with PSO Optimization 55
4.1 Introduction 55
4.2 Introduction of Fuzzy Logic Control 58
4.2.1 Normalization 59
4.2.2 Fuzzification 59
4.2.3 Decision Making Logic 59
4.2.4 Defuzzification 60
4.3 Design of Fuzzy-DDPG 60
4.4 Fuzzy-DDPG Experiments in Simulated Environment 63
4.4.1 Training 64
4.4.2 Testing Results 69
4.5 PSO Trajectory Optimization 71
4.5.1 Particle Initialization 72
4.5.2 Design of Fitness Function 74
4.5.3 Particle Update 79
Chapter 5 Experiments 83
5.1 Introduction 83
5.2 Vision System 85
5.2.1 Object Recognition 86
5.2.2 Object Detection 87
5.2.3 Visual Servo 88
5.2.4 Object Grasp Point 89
5.3 Edge Computer Assisted Computing 89
5.4 Control System 90
5.4.1 End Point Compensation 90
5.5 Experiments 92
5.5.1 Optimized Trajectory Storage 92
5.5.2 Object Grasping by Single Manipulator 94
5.5.3 Object Grasping by Dual Arms 96
5.5.4 Table Cleaning Task 99
5.5.5 Item Packaging Task 105
5.5.6 Flower Arrangement Task 108
Chapter 6 Conclusions and Future Works 112
6.1 Conclusions 112
6.2 Future Works 114
Appendix I 115
Appendix II 117
Appendix III 118
References 120
[1]J. H. Reif, “Complexity of the mover's problem and generalizations, in Proceedings of 20th Annual Symposium on Foundations of Computer Science, 1979, pp. 421-427.
[2]A. A. Hassan, M. El-Habrouk, and S. Deghedie, “Inverse kinematics of redundant manipulators formulated as quadratic programming optimization problem solved using recurrent neural networks: a review, Robotica, pp. 1-18, Nov. 2019.
[3]A. Liegeois, “Automatic supervisory control of the configuration and behavior of multibody mechanisms, IEEE Transactions on Systems, Man, and Cybernetics, vol. 7, no. 12, pp. 868-871, Dec. 1977.
[4]J. Baillieul, J. Hollerbach, and R. Brockett, “Programming and control of kinematically redundant manipulators, in Proceedings of 23rd IEEE Conference on Decision and Control, 1984, pp. 768-774.
[5]J. Baillieul, “Kinematic programming alternatives for redundant manipulators, in Proceedings of 1985 IEEE International Conference on Robotics and Automation, 1985, pp. 722-728.
[6]A. De Luca, L. Lanari, and G. Oriolo, “Control of redundant robots on cyclic trajectories, in Proceedings of 1992 IEEE International Conference on Robotics and Automation, 1992, pp. 500-506.
[7]K. L. Doty, C. Melchiorri, and C. Bonivento, “A theory of generalized inverses applied to robotics, The International Journal of Robotics Research, vol. 12, no. 1, pp. 1-19, Feb. 1993.
[8]R. R. Ulrey, A. A. Maciejewski, and H. J. Siegel, “Parallel algorithms for singular value decomposition, in Proceedings of 8th International Parallel Processing Symposium, 1994, pp. 524-533.
[9]O. Khatib and A. Bowling, “Optimization of the inertial and acceleration characteristics of manipulators, in Proceedings of IEEE International Conference on Robotics and Automation, Minneapolis, MN, USA, 1996, vol. 4, pp. 2883-2889.
[10]M. Goel, A. A. Maciejewski, V. Balakrishnan, and R. W. Proctor, “Failure tolerant teleoperation of a kinematically redundant manipulator: an experimental study, IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans, vol. 33, no. 6, pp. 758-765, Nov. 2003.
[11]C. A. Klein and C. H. Huang, “Review of pseudoinverse control for use with kinematically redundant manipulators, IEEE Transactions on Systems, Man, and Cybernetics, vol. SMC-13, no. 2, pp. 245-250, Mar.-Apr. 1983.
[12]E. S. Conkur, “Path following algorithm for highly redundant manipulator, Robotics and Autonomous Systems, vol. 45, no. 1, pp. 1-22, Oct. 2003.
[13]M. da Graça Marcos, J. A. Tenreiro Machado and T.-P. Azevedo-Perdicoúlis, “A multi-objective approach for the motion planning of redundant manipulators, Applied Soft Computing, vol. 12, no. 2, pp. 589-599, Feb. 2012.
[14]M. M. Abdelhameed and F. A. Tolbah, “A recurrent neural network-based sequential controller for manufacturing automated systems, Mechatronics, vol. 12, no. 4, pp. 617-633, May 2002.
[15]Z. Bingul, H. M. Ertunc, and C. Oysu, “Applying neural network to inverse kinematics problem for 6R robot manipulator with offset wrist, in Proceedings of 7th International Conference on Adaptive and Natural Computing Algorithms, Coimbra, Portugal, 2005, pp. 112-115.
[16]B. Karlik and S. Aydin, “An improved approach to the solution of inverse kinematics problems for robot manipulators, Engineering Applications of Artificial Intelligence, vol. 13, no. 2, pp. 159-164, Apr. 2000.
[17]R. Köker, C. Öz, T. Çakar, and H. Ekiz, “A study of neural network based inverse kinematics solution for a three-joint robot, Robotics and Autonomous Systems, vol. 49, no. 3-4, pp. 227-234, Dec. 2004.
[18]R. Köker, “Reliability-based approach to the inverse kinematics solution of robots using Elman's networks, Engineering Applications of Artificial Intelligence, vo1. 18, no. 6, pp. 685-693, Sep. 2005.
[19]R. Köker, “Design and performance of an intelligent predictive controller for a six-degree-of-freedom robot using the Elman network, Information Sciences, vol. 176, no. 12, pp. 1781-1799, Jun. 2006.
[20]Y. Zhang, Z. Tan, K. Chen, K. Chen, and X. Lv, “Repetitive motion of redundant robots planned by three kinds of recurrent neural networks and illustrated with a four-link planar manipulator’s straight-line example, Robotics and Autonomous Systems, vol. 57, no. 6-7, pp. 645-651, Jun. 2009.
[21]Y. Zhang, X. Lv, Z. Li, Z. Yang, and K. Chen, “Repetitive motion planning of PA10 robot arm subject to joint physical limits and using LVI-based primal-dual neural network, Mechatronics, vol. 18, no. 9, pp. 475-485, Nov. 2008.
[22]Y. Zhang, Z. Tan, Z. Yang, and X. Lv, “A dual neural network applied to drift-free resolution of five-link planar robot arm, in Proceedings of 2008 International Conference on Information and Automation, Zhangjiajie, China, 2008, pp. 1274-1279.
[23]Y. Zhang, H. Zhu, X. Lv, and K. Li, “Joint angle drift problem of PUMA560 robot arm solved by a simplified LVI-based primal-dual neural network, in Proceedings of 2008 IEEE International Conference on Industrial Technology, Chengdu, China, 2008, pp. 1-6.
[24]J. B. Mbede, X. Huang, and M. Wang, “Fuzzy motion planning among dynamic obstacles using artificial potential fields for robot manipulators, Robotics and Autonomous Systems, vol. 32, no. 1, pp. 61-72, Jul. 2000.
[25]C. J. Lin, C. R. Lin S. K. Yu, and C. C. Han, “Singularity avoidance for a redundant robot using fuzzy motion planning, Applied Mechanics and Materials, vol. 479-480, pp. 729-736, Dec. 2013.
[26]J. Chen and H. Y. K. Lau, “A reinforcement motion planning strategy for redundant robot arms based on hierarchical clustering and k-nearest-neighbors, in Proceedings of 2015 IEEE International Conference on Robotics and Biomimetics, Zhuhai, China, 2015, pp. 727-732.
[27]J. K. Parker, A. R. Khoogar, and D. E. Goldberg, “Inverse kinematics of redundant robots using genetic algorithms, in Proceedings of 1989 International Conference on Robotics and Automation, Scottsdale, AZ, USA, 1989, pp. 271-276.
[28]P. Karlra and N. R. Prakash, “A neuro-genetic algorithm approach for solving the inverse kinematics of robotic manipulators, in Proceedings of 2003 IEEE International Conference on Systems, Man and Cybernetics. Conference Theme - System Security and Assurance (Cat. No.03CH37483), Washington, DC, USA, 2003, vol. 2, pp. 1979-1984.
[29]N. Kubota, T. Arakawa, T. Fukuda, and K. Shimojima, “Trajectory generation for redundant manipulator using virus evolutionary genetic algorithm, in Proceedings of International Conference on Robotics and Automation, Albuquerque, NM, USA, 1997, pp. 205-210.
[30]P. Savsani, R. L. Jhala, and V. J. Savsani, “Comparative study of different metaheuristics for the trajectory planning of a robotic arm, IEEE Systems Journal, vol. 10, no. 2, pp. 697-708, Jun. 2016.
[31]T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning, arXiv:1509.02971, 2015.
[32]Introduction to various reinforcement learning algorithms. part I (Q-Learning, SARSA, DQN, DDPG) [Online] Available: https://towardsdatascience.com/introduction-to-various-reinforcement-learning-algorithms-i-q-learning-sarsa-dqn-ddpg-72a5e0cb6287.
[33]H. Najjaran, and A. Goldenberg. “Real-time motion planning of an autonomous mobile manipulator using a fuzzy adaptive kalman filter, Robotics and Autonomous Systems, vol. 55, no. 2, pp. 96-106, Feb. 2007.
[34]J. Jalani and S. Jayaraman, “Design a fuzzy logic controller for a rotary flexible joint robotic arm, in Proceedings of MATEC Web of Conferences, 2018, vol. 150, no.3. Available: https://search.proquest.com/docview/2050502035?accountid=12719. DOI: http://dx.doi.org/10.1051/matecconf/201815001011.
[35]N. Dersarkissian, R. Jia, and D. L. Feitosa, “Control of a two-link robotic arm using fuzzy logic, in Proceedings of 2018 IEEE International Conference on Information and Automation (ICIA), Wuyishan, China, 2018, pp. 481-486.
[36]J. Kennedy and R. Eberhart, “Particle swarm optimization, in Proceedings of ICNN'95 - International Conference on Neural Networks, Perth, WA, Australia, 1995, vol. 4, pp. 1942-1948.
[37]N. Baklouti, H. A. Lamti, K. Salhi, and A. M. Alimi, “PSO based adaptive learning fuzzy logic controller for the irobot create robot, in Proceedings of 13th International Conference on Hybrid Intelligent Systems (HIS 2013), Gammarth, Tunisia, 2013, pp. 99-104.
[38]R. A. Krohling, “Gaussian swarm: a novel particle swarm optimization algorithm, in Proceedings of IEEE Conference on Cybernetics and Intelligent Systems, 2004, vol. 1, pp. 372-376.
[39]A. T. Sadiq, F. A. Raheem, and N. A. F. Abbas, “Optimal trajectory planning of 2-DOF robot arm using the integration of PSO based on D* algorithm and cubic polynomial equation, in Proceedings of the First International Conference for Engineering Researches, 2017, pp. 458-467.
[40]M. Wang, J. Luo, J. Yuan, and U. Walter, “Coordinated trajectory planning of dual-arm space robot using constrained particle swarm optimization, Acta Astronautica, vol. 146, pp. 259-272, May 2018.
[41]Intel RealSense D435i [Online] Available: https://www.intelrealsense.com/depth-camera-d435i/.
[42]LMS100 laser rangefinder [Online] Available:https://www.sick.com/ag/en/detection-and-ranging-solutions/2d-lidar-sensors/lms1xx/c/g91901.
[43]Robotis motors [Online] Available: http://en.robotis.com/.
[44]D-H convention rules [Online] Available: https://blog.robotiq.com/how-to-calculate-a-robots-forward-kinematics-in-5-easy-steps.
[45]Rviz [Online] Available: http://wiki.ros.org/rviz.
[46]Robot operating system [Online] Available: https://www.ros.org/.
[47]URDF file [Online] Available: http://wiki.ros.org/urdf.
[48]V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing atari with deep reinforcement learning, arXiv:1312.5602, 2013.
[49]V. R. Konda and J. N. Tsitsiklis, “Actor-critic algorithms, in Advances in neural information processing systems, 2000, pp. 1008-1014.
[50]V. Mnih, K. Kavukcuoglu, D. Silver, et al., “Human-level control through deep reinforcement learning, Nature, vol. 518, pp. 529-533, 2015.
[51]D. P. Kingma and J. L. Ba, “Adam: a method for stochastic optimization, arXiv:1412.6980, 2014.
[52]I. K. M. Jais, A. R. Ismail, and S. Q. Nisa, “Adam optimization algorithm for wide and deep neural network, Knowledge Engineering and Data Science, vol. 2, no. 1, pp. 41-46, Jun. 2019.
[53]S. Bock and M. Weiß, “A proof of local convergence for the adam optimizer, in Proceedings of 2019 International Joint Conference on Neural Networks (IJCNN), Budapest, Hungary, 2019, pp. 1-8.
[54]S. Ruder, “An overview of gradient descent optimization algorithms, arXiv:1609.04747, 2016.
[55]N. Qian, “On the momentum term in gradient descent learning algorithms, Neural Networks, vol. 12, no. 1, pp. 145-151, Jan. 1999.
[56]J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization, Journal of Machine Learning Research, vol. 12, pp.2121–2159, 2011.
[57]M. D. Zeiler, “ADADELTA: an adaptive learning rate method, arXiv:1212.5701, 2012.
[58]T. Tieleman and G. Hinton, “Lecture 6.5 - RMSProp, COURSERA: Neural Networks for Machine Learning, Technical report, 2012.
[59]M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba, “Hindsight experience replay, in Proceedings of 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 2017, pp. 5055-5065.
[60]D. J. Dubois, Fuzzy Sets and Systems: Theory and Applications, vol. 144, Academic Press, 1980.
[61]L. A. Zadeh, “Fuzzy logic, Computer, vol. 21, no. 4, pp. 83-93, 1988.
[62]D. P. Filev and R. R. Yager, “A generalized defuzzification method via BAD distributions, International Journal of Intelligent Systems, vol. 6, no. 7, pp. 687-697, Dec. 1991.
[63]R. Zhao and R. Govind, “Defuzzification of fuzzy intervals, Fuzzy Sets and Systems, vol. 43, no. 1, pp. 45-55, Sep. 1991.
[64]D. P. Filev and R. R. Yager, “Fuzzy models induced by alternative defuzzification methods, In Proceedings of IEEE 5th International Fuzzy Systems, New Orleans, LA, USA, 1996, vol. 1, pp. 457-462.
[65]C. Edwards and S. Spurgeon, Sliding Mode Control: Theory and Applications, vol. 1, CRC Press, 1998.
[66]Kim and J. Lee, “Trajectory optimization with particle swarm optimization for manipulator motion planning, IEEE Transactions on Industrial Informatics, vol. 11, no. 3, pp. 620-631, Jun. 2015.
[67]K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN, in Proceedings of 2017 IEEE International Conference on Computer Vision (ICCV), 2017.
[68]W. Zhang, Y. Zhang, and N. Liu, “Map-less navigation: a single DRL-based controller for robots with varied dimensions, arXiv:2002.06320, 2020.