Q LEARNING REGRESSION NEURAL NETWORK
DOI:
https://doi.org/10.14311/NNW.2018.%25xAbstract
In this work, a Nadaraya-Watson kerneled learning system which owns
general regression neural network topology is adapted to Q learning method to
evaluate a quick and ecient action selection policy for reinforcement learning
problems. By means of the proposed method Q value function is generalized and
learning speed of Q agent is accelerated. The training data of the developed neural
network are obtained by a standard Q learning agent on closed loop simulation
system. The eciency of the proposed method is tested on populer reinforcement
learning benchmarks and its performance is compared with other popular regression
methods and Q-learning utilized methods.
References
Sutton, Richard S and Barto, Andrew G. Introduction to reinforcement learning. MIT Press, 1998 doi: 10.1007/978-3-319-20010-1_14.
Watkins, Christopher JCH and Dayan, Peter. Q-learning Machine learning. 1992, 8(3-4), pp. 279-292, doi: 10.1007/springerreference_57860.
Sutton, Richard S. Generalization in reinforcement learning: Successful examples using sparse coarse coding Advances in neural information processing systems. 1996, pp. 1038-1044.
Anderson, Charles W. Strategy learning with multilayer connectionist representations Pro-ceedings of the Fourth International Workshop on Machine Learning,1987, pp. 103-114
doi: 10.1016/b978-0-934613-41-5.50014-3
Lin, Long-Ji. Self-improving reactive agents based on reinforcement learning, planning and teaching Machine learning. 1992, 8(3-4), pp. 293-312, doi: 10.1007/978-1-4615-3618-5_5.
Sabes, Philip. Approximating Q-values with basis function representations Proceedings of the Fourth Connectionist Models Summer School, 1993.
Boyan, Justin and Moore, Andrew W. Generalization in reinforcement learning: Safely approximating the value function Advances in neural information processing systems. 1995,
pp. 369-376.
Bradtke, Steven J and Barto, Andrew G. Linear least-squares algorithms for temporal dierence learning Machine learning. 1996, 22(1-3), pp. 33-57, doi: 10.1007/978-0-585-33656-5_4.
Boyan, Justin A. Technical update: Least-squares temporal dierence learning Machine Learning. 2002, 49(2-3), pp. 233{246, doi: 10.1007/978-0-585-33656-5_4.
Brafman, Ronen I and Tennenholtz, Moshe. R-max-a general polynomial time algorithm for near-optimal reinforcement learning Journal of Machine Learning Research. 2002, 3(Oct), pp. 213-231.
Lagoudakis, Michail G and Parr, Ronald. Least-squares policy iteration The Journal of Machine Learning Research. 2003, 4, pp. 1107-1149.
Ernst, Damien and Geurts, Pierre and Wehenkel, Louis. Tree-based batch mode reinforcement learning Journal of Machine Learning Research, 2005, pp. 503-556.
Riedmiller, Martin. Neural tted Q iteration{rst experiences with a data ecient neural reinforcement learning method Machine Learning: ECML 2005, 2005, pp. 317-328. doi: 10.
/11564096_32
Bonarini, Andrea and Lazaric, Alessandro and Restelli, Marcello. Reinforcement learning in complex environments through multiple adaptive partitions Congress of the Italian Association for Articial Intelligence, 2007, pp. 531-542 doi: 10.1007/978-3-540-74782-6_46.
Dutech, Alain and Edmunds, Timothy and Kok, Jelle and Lagoudakis, Michail and Littman, Michael and Riedmiller, Martin and Russell, Bryan and Scherrer, Bruno and Sutton, Richard
and Timmer, Stephan and others. Reinforcement learning benchmarks and bake-os II Advances in Neural Information Processing Systems (NIPS). 2005, 17.
Whiteson, Shimon and Stone, Peter. Evolutionary function approximation for reinforcement learning The Journal of Machine Learning Research, 2006,7, pp. 877-917 doi: 10.1007/
-3-642-13932-1_4.
Cetina, Victor Uc. Multilayer Perceptrons with Radial Basis Functions as Value Functions in Reinforcement Learning. ESANN, 2008, pp. 161-166.
Shibuya, Takeshi and Arita, Hideaki and Hamagami, Tomok. Reinforcement learning in continuous state space with perceptual aliasing by using complex-valued RBF network Systems
Man and Cybernetics (SMC), 2010 IEEE International Conference, 2010, pp. 1799-1803 doi: 10.1109/icsmc.2010.5642294.
Heinen, Milton Roberto and Engel, Paulo Martins. IPNN: An incremental probabilistic neural network for function approximation and regression tasks Neural Networks (SBRN), 2010 Eleventh Brazilian Symposium. 2010, pp. 25{30. doi:10.1109/sbrn.2010.13.
Heinen, Milton Roberto and Engel, Paulo Martins. An incremental probabilistic neural network for regression and reinforcement learning tasks Articial Neural Networks ICANN 2010, 2010, pp. 170{179 doi: 10.1007/978-3-642-15822-3_22.
Heinen, Milton Roberto and Engel, Paulo Martins. IGMN: An incremental connectionist approach for concept formation, reinforcement learning and robotics Journal of Applied
Computing Research, 2011,1(1), pp. 2-19 doi:10.4013/jacr.2011.11.01.
Kobayashi, Takaaki and Shibuya, Takeshi and Morita, Masahiko. Q-learning in continuous state-action space with redundant dimensions by using a selective desensitization neural network Soft Computing and Intelligent Systems (SCIS), 2014 Joint 7th International Conference on and Advanced Intelligent Systems (ISIS), 15th International Symposium. 2014,pp. 801-806. doi: 10.20965/jaciii.2015.p0825.
Specht, Donald F. A general regression neural network Neural Networks, IEEE Transactions on, 1991,2(6), pp. 568-576doi: 10.1016/s0893-6080(09)80013-0.
Closed Loop Simulation System 2013 [accessed 2016-12-06]. Available from: http://ml.informatik.unifreiburg.de/research/clsquare
Wang, Hua O and Tanaka, Kazuo and Grin, Michael F. An approach to fuzzy control of nonlinear systems: stability and design issues Fuzzy Systems, IEEE Transactions on, 1996, 4(1), pp. 14-23 doi: 10.1109/91.481841.
Moore, Andrew William and Hall, Trinity. Ecient memory-based learning for robot control, 1990.
Published
Issue
Section
License
Only original research papers (not published or not simultaneously submitted to another journal) will be reviewed. Before the article is printed, the authors are required to transfer the copyright of the published article to the publisher. The copyright transfer form can be downloaded here: http://www.nnw.cz/copyright.html or here.
Effective upon acceptance for publication in NNW, copyright (including all rights there under and including the right to authorise dissemination, distribution, photocopying and reproduction in all media, whether separately or as a part of a journal issue or otherwise) in the work and any modifications of it by the authors is hereby transferred throughout the world and for the full term and all extensions and renewals thereof, to CTU as a representative of publishers of NNW.
The following rights are retained by the author(s):
- Patent and trademark rights and rights to any process or procedure described in the work.
- The right to photocopy or make single electronic copies of the work for their own personal use, including for their own classroom use, or for the personal use of colleagues, provided that the copies are not offered for sale and are not distributed ina systematic way outside of their employing institution. Posting of a preprint version of the work on an public server is permitted. Posting of the published work on a secure network (not accessible to the public) within the Author's institution is permitted. However, posting of the published work on an electronic public server can only be done with written permission of CTU as a representative of the publishers of NNW.
- The right, subsequent to publication, to use the work or any part there of free of charge in a printed compilation of works of their own, such as collected writings or lecture notes, in a thesis, or to expand the work into book-length form for publication.
All copies, paper or electronic or other use of the work must include an indication of the copyright ownership and a full citation of the journal source.
Abstracting is permitted with credit to the source. For all other copying, reprint or republication permission write to the Czech Technical University in Prague.
Responsibility for the contents of all the publisher papers and letters rests upon the authors and not upon the editors of the NNW.