Meaning guided video captioning

Rushi J. Babariya; Toru Tamaki

arXiv:1912.05730·cs.CV·December 13, 2019·1 cites

Meaning guided video captioning

Rushi J. Babariya, Toru Tamaki

PDF

Open Access 1 Repo

TL;DR

This paper introduces a meaning-guided video captioning model that incorporates object detection and semantic similarity metrics to generate more accurate and meaningful captions, outperforming previous models on the MSDV dataset.

Contribution

It proposes a novel framework that combines object detection with a sequence-to-sequence model and semantic similarity learning for improved video captioning.

Findings

01

Significantly better performance than baseline models.

02

Effective integration of object detection into captioning.

03

Demonstrated on MSDV dataset with improved metrics.

Abstract

Current video captioning approaches often suffer from problems of missing objects in the video to be described, while generating captions semantically similar with ground truth sentences. In this paper, we propose a new approach to video captioning that can describe objects detected by object detection, and generate captions having similar meaning with correct captions. Our model relies on S2VT, a sequence-to-sequence model for video captioning. Given a sequence of video frames, the encoding RNN takes a frame as well as detected objects in the frame in order to incorporate the information of the objects in the scene. The following decoding RNN outputs are then fed into an attention layer and then to a decoder for generating captions. The caption is compared with the ground truth by learning metric so that vector representations of generated captions are semantically similar to those of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

captanlevi/Meaning-guided-video-captioning-
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Video Analysis and Summarization · Advanced Image and Video Retrieval Techniques