Image coding for machines: an end-to-end learned approach

Nam Le; Honglei Zhang; Francesco Cricri; Ramin Ghaznavi-Youvalari; Esa; Rahtu

arXiv:2108.09993·cs.CV·August 31, 2021

Image coding for machines: an end-to-end learned approach

Nam Le, Honglei Zhang, Francesco Cricri, Ramin Ghaznavi-Youvalari, Esa, Rahtu

PDF

TL;DR

This paper introduces the first end-to-end learned image codec optimized for machine consumption, outperforming traditional codecs in object detection and segmentation tasks with significant BD-rate gains.

Contribution

Proposes a neural network-based image codec for machines, with novel training strategies to balance multiple loss functions, achieving superior performance over existing standards.

Findings

01

Outperforms VVC standard on object detection and segmentation

02

Achieves -37.87% and -32.90% BD-rate gains

03

Fast and compact neural network-based codec

Abstract

Over recent years, deep learning-based computer vision systems have been applied to images at an ever-increasing pace, oftentimes representing the only type of consumption for those images. Given the dramatic explosion in the number of images generated per day, a question arises: how much better would an image codec targeting machine-consumption perform against state-of-the-art codecs targeting human-consumption? In this paper, we propose an image codec for machines which is neural network (NN) based and end-to-end learned. In particular, we propose a set of training strategies that address the delicate problem of balancing competing loss functions, such as computer vision task losses, image distortion losses, and rate loss. Our experimental results show that our NN-based codec outperforms the state-of-the-art Versa-tile Video Coding (VVC) standard on the object detection and instance…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.