Q LEARNING REGRESSION NEURAL NETWORK

Authors

DOI:

https://doi.org/10.14311/NNW.2018.%25x

Abstract

In this work, a Nadaraya-Watson kerneled learning system which owns
general regression neural network topology is adapted to Q learning method to
evaluate a quick and ecient action selection policy for reinforcement learning
problems. By means of the proposed method Q value function is generalized and
learning speed of Q agent is accelerated. The training data of the developed neural
network are obtained by a standard Q learning agent on closed loop simulation
system. The eciency of the proposed method is tested on populer reinforcement
learning benchmarks and its performance is compared with other popular regression
methods and Q-learning utilized methods.

Author Biographies

  • Mehmet Sarıgül, Cukurova University
    Computer Engineering Department
  • Mutlu Avcı, Cukurova University
    Biomedical Engineering

References

Sutton, Richard S and Barto, Andrew G. Introduction to reinforcement learning. MIT Press, 1998 doi: 10.1007/978-3-319-20010-1_14.

Watkins, Christopher JCH and Dayan, Peter. Q-learning Machine learning. 1992, 8(3-4), pp. 279-292, doi: 10.1007/springerreference_57860.

Sutton, Richard S. Generalization in reinforcement learning: Successful examples using sparse coarse coding Advances in neural information processing systems. 1996, pp. 1038-1044.

Anderson, Charles W. Strategy learning with multilayer connectionist representations Pro-ceedings of the Fourth International Workshop on Machine Learning,1987, pp. 103-114

doi: 10.1016/b978-0-934613-41-5.50014-3

Lin, Long-Ji. Self-improving reactive agents based on reinforcement learning, planning and teaching Machine learning. 1992, 8(3-4), pp. 293-312, doi: 10.1007/978-1-4615-3618-5_5.

Sabes, Philip. Approximating Q-values with basis function representations Proceedings of the Fourth Connectionist Models Summer School, 1993.

Boyan, Justin and Moore, Andrew W. Generalization in reinforcement learning: Safely approximating the value function Advances in neural information processing systems. 1995,

pp. 369-376.

Bradtke, Steven J and Barto, Andrew G. Linear least-squares algorithms for temporal dierence learning Machine learning. 1996, 22(1-3), pp. 33-57, doi: 10.1007/978-0-585-33656-5_4.

Boyan, Justin A. Technical update: Least-squares temporal dierence learning Machine Learning. 2002, 49(2-3), pp. 233{246, doi: 10.1007/978-0-585-33656-5_4.

Brafman, Ronen I and Tennenholtz, Moshe. R-max-a general polynomial time algorithm for near-optimal reinforcement learning Journal of Machine Learning Research. 2002, 3(Oct), pp. 213-231.

Lagoudakis, Michail G and Parr, Ronald. Least-squares policy iteration The Journal of Machine Learning Research. 2003, 4, pp. 1107-1149.

Ernst, Damien and Geurts, Pierre and Wehenkel, Louis. Tree-based batch mode reinforcement learning Journal of Machine Learning Research, 2005, pp. 503-556.

Riedmiller, Martin. Neural tted Q iteration{rst experiences with a data ecient neural reinforcement learning method Machine Learning: ECML 2005, 2005, pp. 317-328. doi: 10.

/11564096_32

Bonarini, Andrea and Lazaric, Alessandro and Restelli, Marcello. Reinforcement learning in complex environments through multiple adaptive partitions Congress of the Italian Association for Articial Intelligence, 2007, pp. 531-542 doi: 10.1007/978-3-540-74782-6_46.

Dutech, Alain and Edmunds, Timothy and Kok, Jelle and Lagoudakis, Michail and Littman, Michael and Riedmiller, Martin and Russell, Bryan and Scherrer, Bruno and Sutton, Richard

and Timmer, Stephan and others. Reinforcement learning benchmarks and bake-os II Advances in Neural Information Processing Systems (NIPS). 2005, 17.

Whiteson, Shimon and Stone, Peter. Evolutionary function approximation for reinforcement learning The Journal of Machine Learning Research, 2006,7, pp. 877-917 doi: 10.1007/

-3-642-13932-1_4.

Cetina, Victor Uc. Multilayer Perceptrons with Radial Basis Functions as Value Functions in Reinforcement Learning. ESANN, 2008, pp. 161-166.

Shibuya, Takeshi and Arita, Hideaki and Hamagami, Tomok. Reinforcement learning in continuous state space with perceptual aliasing by using complex-valued RBF network Systems

Man and Cybernetics (SMC), 2010 IEEE International Conference, 2010, pp. 1799-1803 doi: 10.1109/icsmc.2010.5642294.

Heinen, Milton Roberto and Engel, Paulo Martins. IPNN: An incremental probabilistic neural network for function approximation and regression tasks Neural Networks (SBRN), 2010 Eleventh Brazilian Symposium. 2010, pp. 25{30. doi:10.1109/sbrn.2010.13.

Heinen, Milton Roberto and Engel, Paulo Martins. An incremental probabilistic neural network for regression and reinforcement learning tasks Articial Neural Networks ICANN 2010, 2010, pp. 170{179 doi: 10.1007/978-3-642-15822-3_22.

Heinen, Milton Roberto and Engel, Paulo Martins. IGMN: An incremental connectionist approach for concept formation, reinforcement learning and robotics Journal of Applied

Computing Research, 2011,1(1), pp. 2-19 doi:10.4013/jacr.2011.11.01.

Kobayashi, Takaaki and Shibuya, Takeshi and Morita, Masahiko. Q-learning in continuous state-action space with redundant dimensions by using a selective desensitization neural network Soft Computing and Intelligent Systems (SCIS), 2014 Joint 7th International Conference on and Advanced Intelligent Systems (ISIS), 15th International Symposium. 2014,pp. 801-806. doi: 10.20965/jaciii.2015.p0825.

Specht, Donald F. A general regression neural network Neural Networks, IEEE Transactions on, 1991,2(6), pp. 568-576doi: 10.1016/s0893-6080(09)80013-0.

Closed Loop Simulation System 2013 [accessed 2016-12-06]. Available from: http://ml.informatik.unifreiburg.de/research/clsquare

Wang, Hua O and Tanaka, Kazuo and Grin, Michael F. An approach to fuzzy control of nonlinear systems: stability and design issues Fuzzy Systems, IEEE Transactions on, 1996, 4(1), pp. 14-23 doi: 10.1109/91.481841.

Moore, Andrew William and Hall, Trinity. Ecient memory-based learning for robot control, 1990.

Additional Files

Published

2018-10-30

Issue

Section

Articles