资源论文Fast Rates for Bandit Optimization with Upper-Confidence Frank-Wolfe

Fast Rates for Bandit Optimization with Upper-Confidence Frank-Wolfe

2020-02-10 | |  96 |   45 |   0

Abstract 

We consider the problem of bandit optimization, inspired by stochastic optimization and online learning problems with bandit feedback. In this problem, the objective is to minimize a global loss function of all the actions, not necessarily a cumulative loss. This framework allows us to study a very general class of problems, with applications in statistics, machine learning, and other fields. To solve this problem, we analyze the Upper-Confidence Frank-Wolfe algorithm, inspired by techniques for bandits and convex optimization. We give theoretical guarantees for the performance of this algorithm over various classes of functions, and discuss the optimality of these results.

上一篇:Stochastic Optimization with Variance Reduction for Infinite Datasets with Finite Sum Structure

下一篇:A General Framework for Robust Interactive Learning

用户评价
全部评价

热门资源

  • The Variational S...

    Unlike traditional images which do not offer in...

  • Learning to Predi...

    Much of model-based reinforcement learning invo...

  • Stratified Strate...

    In this paper we introduce Stratified Strategy ...

  • Learning to learn...

    The move from hand-designed features to learned...

  • A Mathematical Mo...

    Direct democracy, where each voter casts one vo...