资源论文Ob ject Recognition as Machine Translation: Learning a Lexicon for a Fixed Image Vocabulary

Ob ject Recognition as Machine Translation: Learning a Lexicon for a Fixed Image Vocabulary

2020-03-23 | |  67 |   37 |   0

Abstract

We describe a model of ob ject recognition as machine translation. In this model, recognition is a process of annotating image regions with words. Firstly, images are segmented into regions, which are classified into region types using a variety of features. A mapping between region types and keywords supplied with the images, is then learned, using a method based around EM. This process is analogous with learning a lexicon from an aligned bitext. For the implementation we describe, these words are nouns taken from a large vocabulary. On a large test set, the method can predict numerous words with high accuracy. Simple methods identify words that cannot be predicted well. We show how to cluster words that individually are dificult to predict into clusters that can be predicted well — for example, we cannot predict the distinction between train and locomotive using the current set of features, but we can predict the underlying concept. The method is trained on a substantial collection of images. Extensive experimental results illustrate the strengths and weaknesses of the approach.

上一篇:Comparing Intensity Transformations and Their Invariants in the Context of Color Pattern Recognition

下一篇:EigenSegments: A Spatio-Temporal Decomposition of an Ensemble of Images

用户评价
全部评价

热门资源

  • Learning to Predi...

    Much of model-based reinforcement learning invo...

  • Stratified Strate...

    In this paper we introduce Stratified Strategy ...

  • The Variational S...

    Unlike traditional images which do not offer in...

  • A Mathematical Mo...

    Direct democracy, where each voter casts one vo...

  • Rating-Boosted La...

    The performance of a recommendation system reli...