Discovering Latent Discriminative Patterns for Multi-Mode Event Representation

Xie, Wenlong; Yao, Hongxun<sup>*</sup>; Sun, Xiaoshuai; Han, Tingting; Zhao, Sicheng; Chua, Tat-Seng

doi:10.1109/TMM.2018.2879749

摘要

Representation of videos is essential since it conveys an understanding of video content and enables many higher level tasks to be tackled efficiently. However, it is challenging to propose a rational representation for complex event videos, as most video information is either noisy or redundant. In this paper, we propose a compact event representation method that can concisely describe the inner modes of events. We deem that an optimal event representation scheme should reflect the long-term and high-level visual semantics (visual topics) of events, so different from previous frame-level video semantics representation methods and concept-based video representation methods, we investigate the problem from the perspective of segment-level video representations. We then present three appealing properties of segment-level visual semantics. Based on the observation, we propose different algorithms that rely on a novel deep-visual-word-based video encoding method to discover latent discriminative patterns of events. Finally, our multi-mode event representation is obtained by concatenating the discovered patterns as inner modes. We adopt our event representation for representative event parts mining, which can highlight the visual topics of events and remarkably prune the raw videos. We validate our event representation method based on complex event detection task. Experimental results on two standard benchmarking datasets, MED11 and CCV Dataset, show that the proposed method can significantly outperform the state-of-the-art approaches.

出版日期2019-6
单位哈尔滨工业大学

全文

访问全文

收藏分享被引(3) 浏览

更新时间：2024-05-10 16:27

Discovering Latent Discriminative Patterns for Multi-Mode Event Representation

摘要

全文

产品服务

站内浏览

服务支持

联系方式

科研之友