Jointly Learning Environments and Control Policies with Projected Stochastic Gradient Ascent

Journal of Artificial Intelligence Research ◽

10.1613/jair.1.13350 ◽

2022 ◽

Vol 73 ◽

pp. 117-171

Author(s):

Adrien Bolland ◽

Ioannis Boukas ◽

Mathias Berger ◽

Damien Ernst

Keyword(s):

Reinforcement Learning ◽

Time Horizon ◽

Learning Algorithm ◽

Gradient Methods ◽

Optimization Techniques ◽

Small Scale ◽

Joint Design ◽

Gradient Ascent ◽

And Control ◽

Reinforcement Learning Algorithm

We consider the joint design and control of discrete-time stochastic dynamical systems over a finite time horizon. We formulate the problem as a multi-step optimization problem under uncertainty seeking to identify a system design and a control policy that jointly maximize the expected sum of rewards collected over the time horizon considered. The transition function, the reward function and the policy are all parametrized, assumed known and differentiable with respect to their parameters. We then introduce a deep reinforcement learning algorithm combining policy gradient methods with model-based optimization techniques to solve this problem. In essence, our algorithm iteratively approximates the gradient of the expected return via Monte-Carlo sampling and automatic differentiation and takes projected gradient ascent steps in the space of environment and policy parameters. This algorithm is referred to as Direct Environment and Policy Search (DEPS). We assess the performance of our algorithm in three environments concerned with the design and control of a mass-spring-damper system, a small-scale off-grid power system and a drone, respectively. In addition, our algorithm is benchmarked against a state-of-the-art deep reinforcement learning algorithm used to tackle joint design and control problems. We show that DEPS performs at least as well or better in all three environments, consistently yielding solutions with higher returns in fewer iterations. Finally, solutions produced by our algorithm are also compared with solutions produced by an algorithm that does not jointly optimize environment and policy parameters, highlighting the fact that higher returns can be achieved when joint optimization is performed.

Download Full-text

Using Expectation-Maximization for Reinforcement Learning

Neural Computation ◽

10.1162/neco.1997.9.2.271 ◽

1997 ◽

Vol 9 (2) ◽

pp. 271-278 ◽

Cited By ~ 68

Author(s):

Peter Dayan ◽

Geoffrey E. Hinton

Keyword(s):

Reinforcement Learning ◽

Expectation Maximization ◽

Learning Algorithm ◽

Stochastic Gradient ◽

Gradient Ascent ◽

Relative Payoff ◽

The Mean ◽

Stochastic Gradient Ascent ◽

Reinforcement Learning Algorithm

We discuss Hinton's (1989) relative payoff procedure (RPP), a static reinforcement learning algorithm whose foundation is not stochastic gradient ascent. We show circumstances under which applying the RPP is guaranteed to increase the mean return, even though it can make large changes in the values of the parameters. The proof is based on a mapping between the RPP and a form of the expectation-maximization procedure of Dempster, Laird, and Rubin (1977).

Download Full-text

FMR-GA – A Cooperative Multi-agent Reinforcement Learning Algorithm Based on Gradient Ascent

Neural Information Processing - Lecture Notes in Computer Science ◽

10.1007/978-3-319-70087-8_86 ◽

2017 ◽

pp. 840-848 ◽

Cited By ~ 3

Author(s):

Zhen Zhang ◽

Dongqing Wang ◽

Dongbin Zhao ◽

Tingting Song

Keyword(s):

Reinforcement Learning ◽

Learning Algorithm ◽

Gradient Ascent ◽

Multi Agent ◽

Reinforcement Learning Algorithm

Download Full-text

Model dependent reinforcement learning algorithm for reservoir operation stochastic optimization

International Journal of Hydrology ◽

10.15406/ijh.2018.02.00129 ◽

2018 ◽

Vol 2 (5) ◽

Author(s):

Li Wenwu

Keyword(s):

Reinforcement Learning ◽

Stochastic Optimization ◽

Reservoir Operation ◽

Learning Algorithm ◽

Reinforcement Learning Algorithm

Download Full-text

Reinforcement learning algorithm for one-warehouse multi-retailer inventory problem

Automation, Mechanical and Electrical Engineering ◽

10.2495/amee140161 ◽

2014 ◽

Author(s):

C.Y. Li ◽

X.T. Wang ◽

T.W. Zhang

Keyword(s):

Reinforcement Learning ◽

Learning Algorithm ◽

Inventory Problem ◽

Reinforcement Learning Algorithm

Download Full-text

Intelligent Energy Management Strategy Based on an Improved Reinforcement Learning Algorithm With Exploration Factor for a Plug-in PHEV

IEEE Transactions on Intelligent Transportation Systems ◽

10.1109/tits.2021.3085710 ◽

2021 ◽

pp. 1-11

Author(s):

Xinyou Lin ◽

Kuncheng Zhou ◽

Liping Mo ◽

Hailin Li

Keyword(s):

Reinforcement Learning ◽

Energy Management ◽

Management Strategy ◽

Learning Algorithm ◽

Energy Management Strategy ◽

Reinforcement Learning Algorithm

Download Full-text

A multi-objective reinforcement learning algorithm for deadline constrained scientific workflow scheduling in clouds

Frontiers of Computer Science ◽

10.1007/s11704-020-9273-z ◽

2021 ◽

Vol 15 (5) ◽

Author(s):

Yao Qin ◽

Hua Wang ◽

Shanwen Yi ◽

Xiaole Li ◽

Linbo Zhai

Keyword(s):

Reinforcement Learning ◽

Learning Algorithm ◽

Scientific Workflow ◽

Workflow Scheduling ◽

Multi Objective ◽

Reinforcement Learning Algorithm

Download Full-text

Optimization of PV Energy Conversion System Using Reinforcement Learning Algorithm

2020 20th International Conference on Sciences and Techniques of Automatic Control and Computer Engineering (STA) ◽

10.1109/sta50679.2020.9329331 ◽

2020 ◽

Author(s):

Mohamed Ali Zeddini ◽

Mourad Turki ◽

Mohamed Faouzi Mimoun

Keyword(s):

Reinforcement Learning ◽

Energy Conversion ◽

Learning Algorithm ◽

Conversion System ◽

Energy Conversion System ◽

Reinforcement Learning Algorithm

Download Full-text

Enhancing Energy Trading Between Different Islanded Microgrids A Reinforcement Learning Algorithm Case Study in Northern Kordofan State

2020 International Conference on Computer, Control, Electrical, and Electronics Engineering (ICCCEEE) ◽

10.1109/iccceee49695.2021.9429584 ◽

2021 ◽

Author(s):

Moayad ELamin ◽

Fay Elhassan ◽

Mahmoud A. Manzoul

Keyword(s):

Reinforcement Learning ◽

Learning Algorithm ◽

Energy Trading ◽

Reinforcement Learning Algorithm

Download Full-text

Solving flow-shop scheduling problem with a reinforcement learning algorithm that generalizes the value function with neural network

Alexandria Engineering Journal ◽

10.1016/j.aej.2021.01.030 ◽

2021 ◽

Vol 60 (3) ◽

pp. 2787-2800

Author(s):

Jianfeng Ren ◽

Chunming Ye ◽

Feng Yang

Keyword(s):

Neural Network ◽

Reinforcement Learning ◽

Value Function ◽

Flow Shop ◽

Learning Algorithm ◽

Flow Shop Scheduling ◽

Scheduling Problem ◽

Shop Scheduling ◽

The Value Function ◽

Reinforcement Learning Algorithm

Download Full-text

A real-time HIL control system on rotary inverted pendulum hardware platform based on double deep Q-network

Measurement and Control ◽

10.1177/00202940211000380 ◽

2021 ◽

Vol 54 (3-4) ◽

pp. 417-428

Author(s):

Yanyan Dai ◽

KiDong Lee ◽

SukGyu Lee

Keyword(s):

Control System ◽

Reinforcement Learning ◽

Inverted Pendulum ◽

Learning Algorithm ◽

Deep Understanding ◽

Control Engineering ◽

Experience Replay ◽

Real Hardware ◽

Rotary Inverted Pendulum ◽

Reinforcement Learning Algorithm

For real applications, rotary inverted pendulum systems have been known as the basic model in nonlinear control systems. If researchers have no deep understanding of control, it is difficult to control a rotary inverted pendulum platform using classic control engineering models, as shown in section 2.1. Therefore, without classic control theory, this paper controls the platform by training and testing reinforcement learning algorithm. Many recent achievements in reinforcement learning (RL) have become possible, but there is a lack of research to quickly test high-frequency RL algorithms using real hardware environment. In this paper, we propose a real-time Hardware-in-the-loop (HIL) control system to train and test the deep reinforcement learning algorithm from simulation to real hardware implementation. The Double Deep Q-Network (DDQN) with prioritized experience replay reinforcement learning algorithm, without a deep understanding of classical control engineering, is used to implement the agent. For the real experiment, to swing up the rotary inverted pendulum and make the pendulum smoothly move, we define 21 actions to swing up and balance the pendulum. Comparing Deep Q-Network (DQN), the DDQN with prioritized experience replay algorithm removes the overestimate of Q value and decreases the training time. Finally, this paper shows the experiment results with comparisons of classic control theory and different reinforcement learning algorithms.

Download Full-text