• Title of article

    An empirical study of policy convergence in Markov decision process value iteration

  • Author/Authors

    Christopher W. Zobel، نويسنده , , William T. Scherer، نويسنده ,

  • Issue Information
    ماهنامه با شماره پیاپی سال 2005
  • Pages
    16
  • From page
    127
  • To page
    142
  • Abstract
    The value iteration algorithm is a well-known technique for generating solutions to discounted Markov decision process (MDP) models. Although simple to implement, the approach is nevertheless limited in situations where many Markov decision processes must be solved, such as in real-time state-based control problems or in simulation/optimization problems, because of the potentially large number of iterations required for the value function to converge to an -optimal solution. Experimental results suggest, however, that the sequence of solution policies associated with each iteration of the algorithm converges much more rapidly than does the value function. This behavior has significant implications for designing solution approaches for MDPs, yet it has not been explicitly characterized in the literature nor generated significant discussion. This paper seeks to generate such discussion by providing comparative empirical convergence results and exploring several predictors that allow estimation of policy convergence speed based on existing MDP parameters.
  • Keywords
    Markov decision processes , Dynamic programming , convergence results
  • Journal title
    Computers and Operations Research
  • Serial Year
    2005
  • Journal title
    Computers and Operations Research
  • Record number

    928157