ContextVP: Fully Context-Aware Video Prediction

资源分类

2019-10-29 |

146 |

117 |

Abstract. Video prediction models based on convolutional networks, recurrent networks, and their combinations often result in blurry predictions. We identify an important contributing factor for imprecise predictions that has not been studied adequately in the literature: blind spots, i.e., lack of access to all relevant past information for accurately predicting the future. To address this issue, we introduce a fully contextaware architecture that captures the entire available past context for each pixel using Parallel Multi-Dimensional LSTM units and aggregates it using blending units. Our model outperforms a strong baseline network of 20 recurrent convolutional layers and yields state-of-the-art performance for next step prediction on three challenging real-world video datasets: Human 3.6M, Caltech Pedestrian, and UCF-101. Moreover, it does so with fewer parameters than several recently proposed models, and does not rely on deep convolutional networks, multi-scale architectures, separation of background and foreground modeling, motion flflow learning, or adversarial training. These results highlight that full awareness of past context is of crucial importance for video prediction

上一篇：Semantically Aware Urban 3D Reconstruction with Plane-Based Regularization

下一篇：Coloring with Words: Guiding Image Colorization Through Text-based Palette Generation

用户评价

全部评价

还没有评论，说两句吧！

热门资源

Deep Cross-media ...

Cross-media retrieval is a research hotspot in ...
Regularizing RNNs...

Recently, caption generation with an encoder-de...
Learning Expressi...

Facial expression is temporally dynamic event w...
Supervised Descen...

Many computer vision problems (e.
Attributed Graph ...

Graph clustering is a fundamental task which di...

智能在线

400-630-6780
聆听.建议反馈

E-mail: support@tusaishared.com