资源论文A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input

A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input

2020-01-17 | |  134 |   43 |   0

Abstract

We propose a method for automatically answering questions about images by bringing together recent advances from natural language processing and computer vision. We combine discrete reasoning with uncertain predictions by a multiworld approach that represents uncertainty about the perceived world in a bayesian framework. Our approach can handle human questions of high complexity about realistic scenes and replies with range of answer like counts, object classes, instances and lists of them. The system is directly trained from question-answer pairs. We establish a first benchmark for this task that can be seen as a modern attempt at a visual turing test.

上一篇:A Unified Semantic Embedding: Relating Taxonomies and Attributes

下一篇:An Autoencoder Approach to Learning Bilingual Word Representations

用户评价
全部评价

热门资源

  • The Variational S...

    Unlike traditional images which do not offer in...

  • Learning to Predi...

    Much of model-based reinforcement learning invo...

  • Stratified Strate...

    In this paper we introduce Stratified Strategy ...

  • A Mathematical Mo...

    Direct democracy, where each voter casts one vo...

  • Rating-Boosted La...

    The performance of a recommendation system reli...