Learning Environmental Calibration Actions for Policy Self-Evolution?

资源分类

2019-11-05 |

134 |

104 |

Abstract Reinforcement learning in physical world is often expensive. Simulators are commonly employed to train policies. Due to the simulation error, trainedin-simulator policies are hard to be directly deployed in physical world. Therefore, how to efficiently reuse these policies to the real environment is a key issue. To address this issue, this paper presents a policy self-evolution process: in the target environment, the agent firstly executes a few calibration actions to perceive the environment, and then reuses the previous policies according to the observation of the environment. In this way, the mission of policy learning in the target environment is reduced to the task of environment identification through executing the calibration actions, which needs much less samples than learning a policy from scratch. We propose the POSEC (POlicy Self-Evolution by Calibration) approach, which learns the most informative calibration actions for policy self-evolution. Taking three robotic arm controlling tasks as the test beds, we show that the proposed method can learn a fine policy for a new arm with only a few (e.g. five) samples of the target environment.

上一篇：FISH-MML: Fisher-HSIC Multi-View Metric Learning

下一篇：Learning to Design Games: Strategic Environments in Reinforcement Learning

用户评价

全部评价

还没有评论，说两句吧！

热门资源

Deep Cross-media ...

Cross-media retrieval is a research hotspot in ...
Regularizing RNNs...

Recently, caption generation with an encoder-de...
The Variational S...

Unlike traditional images which do not offer in...
Joint Pose and Ex...

Facial expression recognition (FER) is a challe...
Visual Reinforcem...

For an autonomous agent to fulfill a wide range...

智能在线

400-630-6780
聆听.建议反馈

E-mail: support@tusaishared.com