维普中文期刊产品整合服务
1篇 您的检索式:作者名="Douglas Mcilwraith"
    题名 作者 年代 出处 被引量
1Time-in-action RL显示文摘The authors propose a novel reinforcement learning(RL)framework,where agent behaviour is governed by traditional control theory.This integrated approach,called time-in-action RL,enables RL to be applicable to many real-world systems,where underlying dynamics are known in their control theoretical formalism.The key insight to facilitate this integration is to model the explicit time function,mapping the state-action pair to the time accomplishing the action by its underlying controller.In their framework,they describe an action by its value(action value),and the time that it takes to perform(action time).An action-value results from the policy of RL regarding a state.Action time is estimated by an explicit time model learnt from the measured activities of the underlying controller.RL value network is then trained with embedded time model to predict action time.This approach is tested using a variant of Atari Pong and proved to be convergent.Jiangcheng Zhu Zhepei Wang Douglas Mcilwraith Chao Wu Chao Xu Yike Guo 2019IET Cyber-Systems and Robotics2019,1,1:0
返回顶部 每页显示:
共1页 首页 上一页 第1页 下一页 末页 /1 跳转

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费