A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play

David Silver; Thomas Hubert; Julian Schrittwieser; Ioannis Antonoglou; Matthew Lai; Arthur Guez; Marc Lanctot; Laurent Sifre; Dharshan Kumaran; Thore Graepel; Timothy Lillicrap; Karen Simonyan; Demis Hassabis

doi:10.1126/science.aar6404

A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play

Science ◽

10.1126/science.aar6404 ◽

2018 ◽

Vol 362 (6419) ◽

pp. 1140-1144 ◽

Cited By ~ 388

Author(s):

David Silver ◽

Thomas Hubert ◽

Julian Schrittwieser ◽

Ioannis Antonoglou ◽

Matthew Lai ◽

...

Keyword(s):

Artificial Intelligence ◽

Reinforcement Learning ◽

Domain Knowledge ◽

Learning Algorithm ◽

Search Techniques ◽

Domain Specific ◽

Evaluation Functions ◽

History Of ◽

World Champion ◽

Reinforcement Learning Algorithm

The game of chess is the longest-studied domain in the history of artificial intelligence. The strongest programs are based on a combination of sophisticated search techniques, domain-specific adaptations, and handcrafted evaluation functions that have been refined by human experts over several decades. By contrast, the AlphaGo Zero program recently achieved superhuman performance in the game of Go by reinforcement learning from self-play. In this paper, we generalize this approach into a single AlphaZero algorithm that can achieve superhuman performance in many challenging games. Starting from random play and given no domain knowledge except the game rules, AlphaZero convincingly defeated a world champion program in the games of chess and shogi (Japanese chess), as well as Go.

Download Full-text

Research on sports action training method based on generative confrontation network model and artificial intelligence

Journal of Intelligent & Fuzzy Systems ◽

10.3233/jifs-189799 ◽

2021 ◽

pp. 1-11

Author(s):

Yang Yang

Keyword(s):

Artificial Intelligence ◽

Reinforcement Learning ◽

Network Model ◽

Learning Algorithm ◽

Training Data ◽

Generative Adversarial Networks ◽

Movement Training ◽

Model Free ◽

Reward Value ◽

Reinforcement Learning Algorithm

In order to improve the effect of sports movement training, this paper builds a sports movement training model based on artificial intelligence technology based on the generation of confrontation network model. Moreover, in order to achieve the combination of model and model-free deep reinforcement learning algorithm, this paper implements the model’s guidance and constraints on deep reinforcement learning algorithm from the perspective of reward value and behavior strategy and divides the model into two situations. In one case, the existing or manually established expert rules are used as model constraints, which is equivalent to online learning by experts. In another case, expert samples are used as model constraints, and an imitation learning method based on generative adversarial networks is introduced. Moreover, using expert samples as training data, the mechanism that the model is guided by the reward value is combined with the model-free algorithm by generating a confrontation network structure. Finally, this paper studies the performance of the model through experimental research. The research results show that the model constructed in this paper has a certain effect.

Download Full-text

Model dependent reinforcement learning algorithm for reservoir operation stochastic optimization

International Journal of Hydrology ◽

10.15406/ijh.2018.02.00129 ◽

2018 ◽

Vol 2 (5) ◽

Author(s):

Li Wenwu

Keyword(s):

Reinforcement Learning ◽

Stochastic Optimization ◽

Reservoir Operation ◽

Learning Algorithm ◽

Reinforcement Learning Algorithm

Download Full-text

Reinforcement learning algorithm for one-warehouse multi-retailer inventory problem

Automation, Mechanical and Electrical Engineering ◽

10.2495/amee140161 ◽

2014 ◽

Author(s):

C.Y. Li ◽

X.T. Wang ◽

T.W. Zhang

Keyword(s):

Reinforcement Learning ◽

Learning Algorithm ◽

Inventory Problem ◽

Reinforcement Learning Algorithm

Download Full-text

Intelligent Energy Management Strategy Based on an Improved Reinforcement Learning Algorithm With Exploration Factor for a Plug-in PHEV

IEEE Transactions on Intelligent Transportation Systems ◽

10.1109/tits.2021.3085710 ◽

2021 ◽

pp. 1-11

Author(s):

Xinyou Lin ◽

Kuncheng Zhou ◽

Liping Mo ◽

Hailin Li

Keyword(s):

Reinforcement Learning ◽

Energy Management ◽

Management Strategy ◽

Learning Algorithm ◽

Energy Management Strategy ◽

Reinforcement Learning Algorithm

Download Full-text

A multi-objective reinforcement learning algorithm for deadline constrained scientific workflow scheduling in clouds

Frontiers of Computer Science ◽

10.1007/s11704-020-9273-z ◽

2021 ◽

Vol 15 (5) ◽

Author(s):

Yao Qin ◽

Hua Wang ◽

Shanwen Yi ◽

Xiaole Li ◽

Linbo Zhai

Keyword(s):

Reinforcement Learning ◽

Learning Algorithm ◽

Scientific Workflow ◽

Workflow Scheduling ◽

Multi Objective ◽

Reinforcement Learning Algorithm

Download Full-text

Optimization of PV Energy Conversion System Using Reinforcement Learning Algorithm

2020 20th International Conference on Sciences and Techniques of Automatic Control and Computer Engineering (STA) ◽

10.1109/sta50679.2020.9329331 ◽

2020 ◽

Author(s):

Mohamed Ali Zeddini ◽

Mourad Turki ◽

Mohamed Faouzi Mimoun

Keyword(s):

Reinforcement Learning ◽

Energy Conversion ◽

Learning Algorithm ◽

Conversion System ◽

Energy Conversion System ◽

Reinforcement Learning Algorithm

Download Full-text

Enhancing Energy Trading Between Different Islanded Microgrids A Reinforcement Learning Algorithm Case Study in Northern Kordofan State

2020 International Conference on Computer, Control, Electrical, and Electronics Engineering (ICCCEEE) ◽

10.1109/iccceee49695.2021.9429584 ◽

2021 ◽

Author(s):

Moayad ELamin ◽

Fay Elhassan ◽

Mahmoud A. Manzoul

Keyword(s):

Reinforcement Learning ◽

Learning Algorithm ◽

Energy Trading ◽

Reinforcement Learning Algorithm

Download Full-text

Solving flow-shop scheduling problem with a reinforcement learning algorithm that generalizes the value function with neural network

Alexandria Engineering Journal ◽

10.1016/j.aej.2021.01.030 ◽

2021 ◽

Vol 60 (3) ◽

pp. 2787-2800

Author(s):

Jianfeng Ren ◽

Chunming Ye ◽

Feng Yang

Keyword(s):

Neural Network ◽

Reinforcement Learning ◽

Value Function ◽

Flow Shop ◽

Learning Algorithm ◽

Flow Shop Scheduling ◽

Scheduling Problem ◽

Shop Scheduling ◽

The Value Function ◽

Reinforcement Learning Algorithm

Download Full-text

A real-time HIL control system on rotary inverted pendulum hardware platform based on double deep Q-network

Measurement and Control ◽

10.1177/00202940211000380 ◽

2021 ◽

Vol 54 (3-4) ◽

pp. 417-428

Author(s):

Yanyan Dai ◽

KiDong Lee ◽

SukGyu Lee

Keyword(s):

Control System ◽

Reinforcement Learning ◽

Inverted Pendulum ◽

Learning Algorithm ◽

Deep Understanding ◽

Control Engineering ◽

Experience Replay ◽

Real Hardware ◽

Rotary Inverted Pendulum ◽

Reinforcement Learning Algorithm

For real applications, rotary inverted pendulum systems have been known as the basic model in nonlinear control systems. If researchers have no deep understanding of control, it is difficult to control a rotary inverted pendulum platform using classic control engineering models, as shown in section 2.1. Therefore, without classic control theory, this paper controls the platform by training and testing reinforcement learning algorithm. Many recent achievements in reinforcement learning (RL) have become possible, but there is a lack of research to quickly test high-frequency RL algorithms using real hardware environment. In this paper, we propose a real-time Hardware-in-the-loop (HIL) control system to train and test the deep reinforcement learning algorithm from simulation to real hardware implementation. The Double Deep Q-Network (DDQN) with prioritized experience replay reinforcement learning algorithm, without a deep understanding of classical control engineering, is used to implement the agent. For the real experiment, to swing up the rotary inverted pendulum and make the pendulum smoothly move, we define 21 actions to swing up and balance the pendulum. Comparing Deep Q-Network (DQN), the DDQN with prioritized experience replay algorithm removes the overestimate of Q value and decreases the training time. Finally, this paper shows the experiment results with comparisons of classic control theory and different reinforcement learning algorithms.

Download Full-text

A multi-agent reinforcement learning algorithm with fuzzy approximation for Distributed Stochastic Unit Commitment

Journal of Intelligent & Fuzzy Systems ◽

10.3233/jifs-182879 ◽

2019 ◽

Vol 37 (5) ◽

pp. 6613-6628

Author(s):

Ghorbani Farzaneh ◽

Afsharchi Mohsen ◽

Derhami Vali

Keyword(s):

Reinforcement Learning ◽

Unit Commitment ◽

Learning Algorithm ◽

Fuzzy Approximation ◽

Multi Agent ◽

Stochastic Unit Commitment ◽

Reinforcement Learning Algorithm

Download Full-text