A linear programming based approach for composite-action Markov decision processes

Zhang, Zhicong<sup>*</sup>; Li, Shuai; Yan, Xiaohui; Zhang, Liangwei

doi:10.1051/ro/2018081

摘要

We study a time homogeneous discrete composite-action Markov decision process (CMDP) which needs to make multiple decisions at each state. In this particular Markov decision process, the state variables are divided into two separable sets and a two-dimensional composite action is chosen at each decision epoch. To solve a composite-action Markov decision process, we propose a novel linear programming model (Contracted Linear Programming Model, CLPM). We show that the CLPM model obtains the optimal state values of a CMDP process. We analyze and compare the number of variables and constraints of the CLPM model and the Traditional Linear Programming Model (TLPM). Computational experiments compare running times and memory usage of the two models. The CLPM model outperforms the TLPM model in both time complexity and space complexity by theoretical analysis and computational experiments.

出版日期2019-10-9
单位东莞理工学院

全文

访问全文

收藏分享被引浏览

更新时间：2021-06-14 17:28

A linear programming based approach for composite-action Markov decision processes

摘要

全文

产品服务

站内浏览

服务支持

联系方式

科研之友