Reinforcement learning on strategy selection for a cooperative robot system

The high performance and efficiency of multiple unmanned surface vehicles (multi-USV) promote the further civilian and military applications of coordinated USV. As the basis of multiple USVs’ cooperative work, considerable attention has been spent on developing the decentralized formation control of the USV swarm. Formation control of multiple USV belongs to the geometric problems of a multi-robot system. The main challenge is the way to generate and maintain the formation of a multi-robot system. The rapid development of reinforcement learning provides us with a new solution to deal with these problems. In this paper, we introduce a decentralized structure of the multi-USV system and employ reinforcement learning to deal with the formation control of a multi-USV system in a leader–follower topology. Therefore, we propose an asynchronous decentralized formation control scheme based on reinforcement learning for multiple USVs. First, a simplified USV model is established. Simultaneously, the formation shape model is built to provide formation parameters and to describe the physical relationship between USVs. Second, the advantage deep deterministic policy gradient algorithm (ADDPG) is proposed. Third, formation generation policies and formation maintenance policies based on the ADDPG are proposed to form and maintain the given geometry structure of the team of USVs during movement. Moreover, three new reward functions are designed and utilized to promote policy learning. Finally, various experiments are conducted to validate the performance of the proposed formation control scheme. Simulation results and contrast experiments demonstrate the efficiency and stability of the formation control scheme.

Download Full-text

Using game theory to describe strategy selection for environmental risk and carbon emissions reduction in the green supply chain

Journal of Loss Prevention in the Process Industries ◽

10.1016/j.jlp.2012.05.004 ◽

2012 ◽

Vol 25 (6) ◽

pp. 927-936 ◽

Cited By ~ 94

Author(s):

Rui Zhao ◽

Gareth Neighbour ◽

Jiaojie Han ◽

Michael McGuire ◽

Pauline Deutz

Keyword(s):

Game Theory ◽

Supply Chain ◽

Environmental Risk ◽

Carbon Emissions ◽

Green Supply Chain ◽

Emissions Reduction ◽

Strategy Selection ◽

Carbon Emissions Reduction ◽

Selection For

Download Full-text

Hypothesis Testing: Strategy Selection for Generalising versus Limiting Hypotheses

Thinking & Reasoning ◽

10.1080/135467899394084 ◽

1999 ◽

Vol 5 (1) ◽

pp. 67-92 ◽

Cited By ~ 11

Author(s):

Barbara A. Spellman

Keyword(s):

Hypothesis Testing ◽

Strategy Selection ◽

Testing Strategy ◽

Selection For

Download Full-text

Adaptive motion selection for online hand–eye calibration

Robotica ◽

10.1017/s0263574707003426 ◽

2007 ◽

Vol 25 (5) ◽

pp. 529-536

Author(s):

Jing Zhang ◽

Fanhuai Shi ◽

Yuncai Liu

Keyword(s):

Motion Planning ◽

Dynamic Threshold ◽

Selection Algorithm ◽

Robot System ◽

End Effector ◽

Threshold Determination ◽

Robot Gripper ◽

Adaptive Motion ◽

Selection For ◽

Relative Pose

SUMMARYWhile a robot moves, online hand–eye calibration to determine the relative pose between the robot gripper/end-effector and the sensors mounted on it is very important in a vision-guided robot system. During online hand–eye calibration, it is impossible to perform motion planning to avoid degenerate motions and small rotations, which may lead to unreliable calibration results. This paper proposes an adaptive motion selection algorithm for online hand–eye calibration, featured by dynamic threshold determination for motion selection and getting reliable hand–eye calibration results. Simulation and real experiments demonstrate the effectiveness of our method.

Download Full-text