资源论文logistic markov decision processes

logistic markov decision processes

2019-10-31 | |  52 |   26 |   0
Abstract User modeling in advertising and recommendation has typically focused on myopic predictors of user responses. In this work, we consider the long-term decision problem associated with user interaction. We propose a concise specification of long-term interaction dynamics by combining factored dynamic Bayesian networks with logistic predictors of user responses, allowing state-of-the-a prediction models to be seamlessly extended. We show how to solve such models at scale by providing a constraint generation approach for approximate linear programming that overcomes the variable coupling and nonlinearity induced by the logistic regression predictor. T efficacy of the approach is demonstrated on advertising domains with up to 254 states and 239 actions.

上一篇:weighted double q learning

下一篇:microblog sentiment classi cation via recurrent random walk network learning

用户评价
全部评价

热门资源

  • Learning to Predi...

    Much of model-based reinforcement learning invo...

  • Stratified Strate...

    In this paper we introduce Stratified Strategy ...

  • The Variational S...

    Unlike traditional images which do not offer in...

  • A Mathematical Mo...

    Direct democracy, where each voter casts one vo...

  • Rating-Boosted La...

    The performance of a recommendation system reli...