Look and Modify: Modification Networks for Image Captioning

Fawaz Sammani; Mahmoud Elsayed

arXiv:1909.03169·cs.CV·March 10, 2020·1 cites

Look and Modify: Modification Networks for Image Captioning

Fawaz Sammani, Mahmoud Elsayed

PDF

Open Access 1 Repo

TL;DR

This paper proposes a novel modification network for image captioning that learns to refine existing captions by focusing on residual information, leading to improved caption quality on the COCO dataset.

Contribution

It introduces a new framework that modifies captions by modeling residual information, enhancing existing captioning models without generating captions from scratch.

Findings

01

Improved caption quality on COCO dataset.

02

Effective modification of captions across multiple frameworks.

Abstract

Attention-based neural encoder-decoder frameworks have been widely used for image captioning. Many of these frameworks deploy their full focus on generating the caption from scratch by relying solely on the image features or the object detection regional features. In this paper, we introduce a novel framework that learns to modify existing captions from a given framework by modeling the residual information, where at each timestep the model learns what to keep, remove or add to the existing caption allowing the model to fully focus on "what to modify" rather than on "what to predict". We evaluate our method on the COCO dataset, trained on top of several image captioning frameworks and show that our model successfully modifies captions yielding better ones with better evaluation scores.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

fawazsammani/look-and-modify
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Human Pose and Action Recognition · Domain Adaptation and Few-Shot Learning