Boosting Video Captioning with Dynamic Loss Network

Nasib Ullah; Partha Pratim Mohanta

arXiv:2107.11707·cs.CV·February 3, 2022

Boosting Video Captioning with Dynamic Loss Network

Nasib Ullah, Partha Pratim Mohanta

PDF

Open Access

TL;DR

This paper introduces a dynamic loss network for video captioning that directly optimizes evaluation metrics, leading to improved performance over existing methods on standard datasets.

Contribution

The paper proposes a novel dynamic loss network that provides direct feedback based on evaluation metrics, enhancing video captioning accuracy.

Findings

01

Outperforms previous methods on MSVD and MSRVTT datasets.

02

Provides more efficient optimization compared to reinforcement learning approaches.

03

Easily adaptable to similar vision-language tasks.

Abstract

Video captioning is one of the challenging problems at the intersection of vision and language, having many real-life applications in video retrieval, video surveillance, assisting visually challenged people, Human-machine interface, and many more. Recent deep learning based methods have shown promising results but are still on the lower side than other vision tasks (such as image classification, object detection). A significant drawback with existing video captioning methods is that they are optimized over cross-entropy loss function, which is uncorrelated to the de facto evaluation metrics (BLEU, METEOR, CIDER, ROUGE). In other words, cross-entropy is not a proper surrogate of the true loss function for video captioning. To mitigate this, methods like REINFORCE, Actor-Critic, and Minimum Risk Training (MRT) have been applied but have limitations and are not very effective. This paper…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Human Pose and Action Recognition · Advanced Image and Video Retrieval Techniques

MethodsREINFORCE