Temporal Action Localization by Structured Maximal Sums

资源分类

2019-12-04 |

41 |

29 |

Abstract We address the problem of temporal action localization in videos. We pose action localization as a structured prediction over arbitrary-length temporal windows, where each window is scored as the sum of frame-wise classifi- cation scores. Additionally, our model classifies the start, middle, and end of each action as separate components, allowing our system to explicitly model each action’s temporal evolution and take advantage of informative temporal dependencies present in this structure. In this framework, we localize actions by searching for the structured maximal sum, a problem for which we develop a novel, provablyefficient algorithmic solution. The frame-wise classification scores are computed using features from a deep Convolutional Neural Network (CNN), which are trained end-toend to directly optimize for a novel structured objective. We evaluate our system on the THUMOS ’14 action detection benchmark and achieve competitive performance.

上一篇：Template Matching with Deformable Diversity Similarity

下一篇：Temporal Convolutional Networks for Action Segmentation and Detection

用户评价

全部评价

还没有评论，说两句吧！

热门资源

Learning to Predi...

Much of model-based reinforcement learning invo...
Stratified Strate...

In this paper we introduce Stratified Strategy ...
The Variational S...

Unlike traditional images which do not offer in...
A Mathematical Mo...

Direct democracy, where each voter casts one vo...
Rating-Boosted La...

The performance of a recommendation system reli...

智能在线

400-630-6780
聆听.建议反馈

E-mail: support@tusaishared.com