Object Level Visual Reasoning in Videos
Fabien Baradel, Natalia Neverova, Christian Wolf, Julien Mille, Greg, Mori

TL;DR
This paper introduces a novel object-level reasoning model for video activity recognition that captures detailed semantic interactions between objects and actors, achieving state-of-the-art results on multiple datasets.
Contribution
The paper presents a new model that performs semantic spatiotemporal reasoning at the object level using advanced detection networks, enhancing activity understanding.
Findings
Achieves state-of-the-art performance on three standard datasets.
Effectively models detailed object interactions in videos.
Provides visualizations of learned semantic interactions.
Abstract
Human activity recognition is typically addressed by detecting key concepts like global and local motion, features related to object classes present in the scene, as well as features related to the global context. The next open challenges in activity recognition require a level of understanding that pushes beyond this and call for models with capabilities for fine distinction and detailed comprehension of interactions between actors and objects in a scene. We propose a model capable of learning to reason about semantically meaningful spatiotemporal interactions in videos. The key to our approach is a choice of performing this reasoning at the object level through the integration of state of the art object detection networks. This allows the model to learn detailed spatial interactions that exist at a semantic, object-interaction relevant level. We evaluate our method on three standard…
Click any figure to enlarge with its caption.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7|
mAP |
bag |
bed |
bedding |
book/papers |
bottle/tube |
bowl |
box |
brush |
cabinet |
cell-phone |
clothing |
cup |
door |
drawers |
food |
fork |
knife |
laptop |
microwave |
oven |
pen/pencil |
pillow |
plate |
refrigerator |
sink |
spoon |
stuffed animal |
table |
toothbrush |
towel |
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| R50 [16] | 40.5 | 29.7 | 68.9 | 65.8 | 64.5 | 58.2 | 33.1 | 22.1 | 19.0 | 23.9 | 54.0 | 45.5 | 28.6 | 49.2 | 28.7 | 49.6 | 19.4 | 37.5 | 62.9 | 48.8 | 23.0 | 36.9 | 39.2 | 12.5 | 55.9 | 58.8 | 31.1 | 57.4 | 26.8 | 39.6 | 22.9 |
| I3D [6] | 39.7 | 24.9 | 71.7 | 71.4 | 62.5 | 57.1 | 27.1 | 19.2 | 33.9 | 20.7 | 50.6 | 45.8 | 24.7 | 54.7 | 19.1 | 50.8 | 19.3 | 41.9 | 54.0 | 27.5 | 21.4 | 37.4 | 42.9 | 12.6 | 42.5 | 60.4 | 33.9 | 46.0 | 23.5 | 59.6 | 34.7 |
| Ours | 44.7 | 30.2 | 72.3 | 70.7 | 64.9 | 59.8 | 38.2 | 24.6 | 26.3 | 22.4 | 64.5 | 47.2 | 35.4 | 57.9 | 25.2 | 48.5 | 24.5 | 40.2 | 72.0 | 54.1 | 26.5 | 39.9 | 48.6 | 15.2 | 53.5 | 60.7 | 36.8 | 52.8 | 27.9 | 64.0 | 37.6 |
| Method | Object type | EPIC | VLOG | SS | |||
|---|---|---|---|---|---|---|---|
| obj. head | 2 heads | obj. head | 2 heads | obj. head | 2 heads | ||
| Baseline | - | - | 38.33 | - | 35.03 | - | 31.31 |
| ORN | pixel | 23.71 | 38.83 | 14.40 | 35.18 | 2.51 | 31.43 |
| ORN | COCO | 29.94 | 40.89 | 27.14 | 37.49 | 10.26 | 32.12 |
| ORN-mlp | COCO | 28.15 | 39.41 | 25.40 | 36.35 | - | - |
| ORN | COCO-visual | 28.45 | 38.92 | 22.92 | 35.49 | - | - |
| ORN | COCO-shape | 21.92 | 37.16 | 7.18 | 35.39 | - | - |
| ORN | COCO-class | 21.96 | 37.75 | 13.40 | 35.94 | - | - |
| ORN | COCO-intra | 29.25 | 38.10 | 26.78 | 36.28 | - | - |
| ORN clique-1 | COCO | 28.25 | 40.18 | 26.48 | 36.71 | - | - |
| ORN clique-3 | COCO | 22.61 | 37.67 | 27.05 | 36.04 | - | - |
| Conv1 | Conv2 | Conv3 | Conv4 | Conv5 | Aggreg | SS | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2D | 3D | 2.5D | 2D | 3D | 2.5D | 2D | 3D | 2.5D | 2D | 3D | 2.5D | 2D | 3D | 2.5D | GAP | RNN | |
| ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | 15.73 |
| ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | - | ✓ | 15.88 |
| - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | ✓ | - | 31.42 |
| - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | ✓ | - | 27.58 |
| ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | - | ✓ | - | ✓ | - | 31.28 |
| ✓ | - | - | ✓ | - | - | ✓ | - | - | - | ✓ | - | - | ✓ | - | ✓ | - | 32.06 |
| ✓ | - | - | ✓ | - | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | ✓ | - | 32.25 |
| ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | - | - | ✓ | ✓ | - | 31.31 |
| ✓ | - | - | ✓ | - | - | ✓ | - | - | - | - | ✓ | - | - | ✓ | ✓ | - | 32.79 |
| ✓ | - | - | ✓ | - | - | - | - | ✓ | - | - | ✓ | - | - | ✓ | ✓ | - | 33.77 |
| - | ✓ | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | 28.71 |
| - | ✓ | - | - | ✓ | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | 31.42 |
| - | - | ✓ | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | 20.05 |
| - | - | ✓ | - | - | ✓ | ✓ | - | - | ✓ | - | - | ✓ | - | - | ✓ | - | 22.52 |
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsHuman Pose and Action Recognition · Multimodal Machine Learning Applications · Video Analysis and Summarization
11institutetext: Université Lyon, INSA Lyon, CNRS, LIRIS, F-69621, Villeurbanne, France, 11email: [email protected] 22institutetext: Facebook AI Research, Paris, France, 22email: [email protected] 33institutetext: INRIA, CITI Laboratory, Villeurbanne, France 44institutetext: Laboratoire d’Informatique de l’Univ. de Tours, INSA Centre Val de Loire,
41034, Blois, France, 44email: [email protected] 55institutetext: Simon Fraser University, Vancouver, Canada, 55email: [email protected]
https://fabienbaradel.github.io/eccv18_object_level_visual_reasoning/
Object Level Visual Reasoning in Videos
Fabien Baradel 11
Natalia Neverova 22
Christian Wolf 1133
Julien Mille 44
Greg Mori 55
Abstract
Human activity recognition is typically addressed by detecting key concepts like global and local motion, features related to object classes present in the scene, as well as features related to the global context. The next open challenges in activity recognition require a level of understanding that pushes beyond this and call for models with capabilities for fine distinction and detailed comprehension of interactions between actors and objects in a scene. We propose a model capable of learning to reason about semantically meaningful spatio-temporal interactions in videos. The key to our approach is a choice of performing this reasoning at the object level through the integration of state of the art object detection networks. This allows the model to learn detailed spatial interactions that exist at a semantic, object-interaction relevant level. We evaluate our method on three standard datasets (Twenty-BN Something-Something, VLOG and EPIC Kitchens) and achieve state of the art results on all of them. Finally, we show visualizations of the interactions learned by the model, which illustrate object classes and their interactions corresponding to different activity classes.
Keywords:
Video understanding Human-object interaction
1 Introduction
The field of video understanding is extremely diverse, ranging from extracting highly detailed information captured by specifically designed motion capture systems [33] to making general sense of videos from the Web [1]. As in the domain of image recognition, there exist a number of large-scale video datasets [6, 26, 13, 12, 23, 14], which allow the training of high-capacity deep learning models from massive amounts of data. These models enable detection of key cues present in videos, such as global and local motion, various object categories and global scene-level information, and often achieve impressive performance in recognizing high-level, abstract concepts in the wild.
However, recent attention has been directed toward a more thorough understanding of human-focused activity in diverse internet videos. These efforts range from atomic human actions [14] to fine-grained object interactions [13] to everyday, commonly occurring human-object interactions [12]. This returns us to a human-centric viewpoint of activity recognition where it is not only the presence of certain objects / scenes that dictate the activity present, but the manner, order, and effects of human interaction with these scene elements that are necessary for understanding. In a sense, this is akin to the problems in current 3D human activity recognition datasets [33], but requires the more challenging reasoning and understanding of diverse environments common to internet video collections.
Humans are able to infer what happened in a video given only a few sample frames. In particular, they can infer complex activities happening between pairs of frames. This faculty is called reasoning and is a key component of human intelligence. As an example we can consider the pair of images in Figure 1, which shows a complex situation involving articulated objects (human, carrots and knife), the change of location and composition of objects. For humans it is straightforward to draw a conclusion on what happened (a carrot was chopped by the human). Humans have this extraordinary ability of performing visual reasoning on very complicated tasks while it remains unattainable for contemporary computer vision algorithms [37, 11].
The ability to perform visual reasoning in computer vision algorithms is still an open problem. Attempts have been made for learning interactions between different entities in images with promising results on Visual Question Answering. There have been a number of attempts to equip neural models with reasoning abilities by training them to solve Visual Question Answering (VQA) problems. Among proposed solutions are prior-less data normalization [27], structuring networks to model relationships [32, 43] as well as more complex attention based mechanisms [18]. At the same time, it was shown that high performance on existing VQA datasets can be achieved by simply discovering biases in the data [21].
We extend these efforts to object level reasoning in videos. Since a video is a temporal sequence, we leverage time as an explicit causal signal to identify causal object relations. Our approach is related to the concept of the “arrow of the time” [28] involving the “one-way direction” or “asymmetry” of time. Causal event occurs before the event it affects (). In Figure 1 the knife was used before the carrot switched over to the chopped-up state on the right side. For a video classification problem, we want to identify a causal event happening in a video that affects its label . But instead of identifying this causal event directly from pixels we want to identify it from an object level perspective. We believe that such an approach would be able to learn causal signals.
Following this hypothesis we propose to make a bridge between object detection and activity recognition. Object detection allows us to extract low-level information from a scene with all the present object instances and their semantic meanings. However, detailed activity understanding requires reasoning over these semantic structures, determining which objects were involved in interactions, of what nature, and what were the results of these. To compound problems, the semantic structure of a scene may change during a video (e.g. a new object can appear, a person may make a move from one point to another one of the scene).
We propose an Object Relation Network (ORN), a neural network module for reasoning between detected semantic object instances through space and time. The ORN has potential to address these issues and conduct relational reasoning over object interactions for the purpose of activity recognition. A set of object detection masks ranging over different object categories and temporal occurrences is input to the ORN. The ORN is able to infer pairwise relationships between objects detected at varying different moments in time.
Code and object masks predictions will be publicly available111https://github.com/fabienbaradel/object_level_visual_reasoning.
2 Related work
Action Recognition. Action recognition has a long history in computer vision. Pre-deep learning approaches in action recognition focused on handcrafted spatio-temporal features including space-time interest points like SIFT-3D, HOG3D, IDT and aggregated them using bag-of-words techniques. Some hand-crafted representations, like dense trajectories [42], still give competitive performance and are frequently combined with deep learning.
In the recent past, work has shifted to deep learning. Early attempts adapt 2D convolutional networks to videos through temporal pooling and 3D convolutions [2, 40]. 3D convolutions are now widely adopted for activity recognition with the introduction of feature transfer by inflating pre-trained 2D convolutional kernels from image classification models trained on ImageNet/ILSVRC [31] through 3D kernels [6]. The downside of 3D kernels is their computational complexity and the large number of learnable parameters, leading to the introduction of 2.5D kernels, i.e. separable filters in the form of a 2D spatial kernel followed by a temporal kernel [44]. An alternative to temporal convolutions are Recurrent Neural Networks (RNNs) in their various gated forms (GRUs, LSTMs) [17, 9].
Karpathy et al. [20] presented a wide study on different ways of connecting information in spatial and temporal dimensions through convolutions and pooling. On very general datasets with coarse activity classes they have showed that there was a small margin between classifying individual frames and classifying videos with more sophisticated temporal aggregation.
Simoyan et al. [35] proposed a widely adopted two-stream architecture for action recognition which extracts two different streams, one processing raw RGB input and one processing pre-computed optical flow images. The method outperformed the state of the art, but relies on rather small scale optical flow computations.
In slightly narrower settings, prior information on the video content can allow more fine-grained models. Articulated pose is widely used in cases where humans are guaranteed to be present [33]. Pose estimation and activity recognition as a joint (multi-task) problem has recently shown to improve both tasks [25]. Somewhat related to our work, Structural RNNs [19] perform activity recognition by integrating features from semantic objects and their relationships. However, they handle the temporal evolution of tracked objects in videos with a set of RNNs, each of which corresponds to cliques in a graph which models the spatio-temporal relationships between these objects. This graph is hand-crafted manually for each application, though related work provides learnable connections via gating functions [8].
Attention models are a way to structure deep networks in an often generic way. They are able to iteratively focus attention to specific parts in the data without requiring prior knowledge about part or object positions. In activity recognition, they have gained some traction in recent years, either as soft-attention on articulated pose (joints) [36], on feature map cells [34, 39], on time [45] or on parts in raw RGB input through differentiable crops [3].
When raw video data is globally fed into deep neural networks, they focus on extracting spatio-temporal features and perform aggregations. It has been shown that these techniques fail on challenging fine-grained datasets, which require learning long temporal dependencies and human-object interactions. A concentrated effort has been made to create large scale datasets to overcome these issues [13, 12, 23, 14].
Relational Reasoning. Relational reasoning is a well studied field for many applications ranging from visual reasoning [32] to reasoning about physical systems [4]. Battaglia et al. [4] introduce a fully-differentiable network physics engine called Interaction Network (IN). IN learns to predict several physical systems such as gravitational systems, rigid body dynamics, and mass-spring systems. It shows impressive results; however, it learns from a virtual environment, which provides access to virtually unlimited training examples. Following the same perspective, Santoro et al. [32] introduced Relation Network (RN), a plug-in module for reasoning in deep networks. RN shows human-level performance in Visual Question Answering (VQA) by inferring pairwise “object” relations. However, in contrast to our work, the term “object” in [32] does not refer to semantically meaningful entities, but to discrete cells in feature maps. The number of interactions therefore grows with feature map resolutions, which makes it difficult to scale. Furthermore, a recent study [21] has shown that some of these results are subject to dataset bias and do not generalize well to small changes in the settings of the dataset.
In the same line, a recent work [38] has shown promising results on discovering objects and their interactions in an unsupervised manner using training examples from virtual environments. In [41], attention and relational modules are combined on a graph structure. From a different perspective, [27] show that relational reasoning can be learned for visual reasoning in a data driven way without any prior using conditional batch normalization with a feature-wise affine transformation based on conditioning information. In an opposite approach, a strong structural prior is learned in the form of a complex attention mechanism: in [18], an external memory module combined with attention processes over input images and text questions, performing iterative reasoning for VQA.
While most of the discussed work has been designed for VQA and for predictions on physical systems and environments, extensions have been proposed for video understanding. Reasoning in videos on a mask or segmentation level has been attempted for video prediction [24], where the goal was to leverage semantic information to be able predict further into the future. Zhou et al [5] have recently shown state-of-the-art performance on challenging datasets by extending Relation Network to video classification. Their chosen entities are frames, on which they employ RN to reason on a temporal level only through pairwise frame relations. The approach is promising, but restricted to temporal contextual information without an understanding on a local object level, which is provided by our approach.
Reasoning over sets of objects is somewhat related to reasoning from unstructured data points, as done in PointNet [29], designed to learn from unordered sets of points. PointNet shares many properties with DeepSet [46] which is a more general framework for extracting information from sets of objects. To some extent, our work is related to PointNet, as we handle unordered sets of objects in a permutation invariant way. However, we have an object relation viewpoint that directly reasons over relationships between these semantic entities.
3 Object-level Visual Reasoning in Space and Time
Our goal is to extract multiple types of cues from a video sequence: interactions between predicted objects and their semantic classes, as well as local and global motion in the scene. We formulate this objective as a neural architecture with two heads: an activity head and an object head. Figure 2 gives a functional overview of the model. Both heads share common features up to a certain layer shown in red in the figure. The activity head, shown in orange in the figure, is a CNN-based architecture employing convolutional layers, including spatio-temporal convolutions, able to extract global motion features. However, it is not able to extract information from an object level perspective. We leverage the object head to perform reasoning on the relationships between predicted object instances.
Our main contribution is a new structured module called Object Relation Network (ORN), which is able to perform spatio-temporal reasoning between detected object instances in the video. ORN is able to reason by modeling how objects move, appear and disappear and how they interact between two frames.
In this section, we will first describe our main contribution, the ORN network. We then provide details about object instance features, about the activity head, and finally about the final recognition task. In what follows, lowercase letters denote 1D vectors while uppercase letters are used for 2D and 3D matrices or higher order tensors. We assume that the input of our system is a video of frames denoted by \mbox{\mathbf{X}}_{1:T}=(\mbox{\mathbf{X}}_{t})_{t=1}^{T} where \mbox{\mathbf{X}}_{t} is the RGB image at timestep . The goal is to learn a mapping from \mbox{\mathbf{X}}_{1:T} to activity classes .
3.1 Object Relation Network
ORN (Object Relation Network) is a module for reasoning between semantic objects through space and time. It captures object moves, arrivals and interactions in an efficient manner. We suppose that for each frame , we have a set of objects with associated features \mbox{\mathbf{o}}_{t}^{k}. Objects and features are detected and computed by the object head described in Section 3.2.
Reasoning about activities in videos is inherently temporal, as activities follow the arrow of time [28], i.e. the causality of the time dimension imposes that past actions have consequences in the future but not vice-versa. We handle this by sampling: running a process over time , and for each instant , sampling a second frame with . Our network reasons on objects which interact between pairs of frames and their corresponding sets of objects \mbox{\mathbf{O}}_{t^{\prime}}=\big{\{}\mbox{\mathbf{o}}^{k}_{t^{\prime}}\big{\}}^{K^{\prime}}_{k=1} and \mbox{\mathbf{O}}_{t}=\big{\{}\mbox{\mathbf{o}}^{k}_{t}\big{\}}^{K}_{k=1}. The goal is to learn a general function defined on the set of all input objects from the combined set of both frames:
[TABLE]
The objects in this set are unordered, aside for the frame they belong to.
This task is related to a problem raised in the PointNet algorithm [29] discussed in Section 2. PointNet approximates a general function over an item set as a symmetric function on transformed elements of the set
[TABLE]
In [29], is a max pooling operation. The authors show that it allows universal approximation of continuous sets of functions given that the hidden representation (the output of the mapping ) is of sufficiently high dimension.
We argue that the approximation in (2) can be extended as follows:
[TABLE]
where is the set of cliques of a graph defined over the item set , is the concatenation operator and we chose the sum operator as symmetric function. The input dimension of the non-linearity depends on the size of clique but maps to a fixed output dimension . In the case where is composed of unary cliques only, form (3) decomposes like (2) with the exception of a different symmetry operator (sum instead of max pooling). Choosing different graphical structures through will lead to different terms in the summation and allows modeling different types of interactions between items in the item set.
Note that the interactions between items in (3) are not exclusively modeled through . Indeed, it is interesting to note, that the graphical decomposition provided by leads to interactions which are different from the interactions the same decomposition would provide when used in a probabilistic graphical model, like for instance a Markov Random Field (MRF). In particular, a decomposition into unary terms only, as given in equation (2), does not lead to independence between items, whereas an MRF with unary terms only is equivalent to a distribution over independent random variables. This is a consequence of the global mapping , which is defined on the sum over all direct interactions. Higher-order interactions between several items not directly modeled through a non-linearity can eventually be learned by the model through the joint output space of all , provided that the dimensionality of this space is high enough to incorporate all interactions. However, whereas the mapping provides a direct model of interactions between pairs of items, learning interactions between two items , which are not directly captured through a clique and its corresponding , requires learning a corresponding subspace in the common output space spanned by all .
This leads to the question of how to define the trade-off between the complexity of the decomposition and the output dimensionality of the mapping , both of which will determine the complexity of the modeled interactions. Increasing the size of cliques in will increase the input dimension (and therefore the capacity) of the mapping as well as the computational complexity of the sum operation.
Inspired by relational networks [32], we chose to directly model inter-frame interactions between pairs of objects and leave modeling of higher-order interactions to the output space of the mappings and the global mapping :
[TABLE]
In order to better directly model long-range interactions, we make the global mapping recurrent, which leads to the following form:
[TABLE]
where \mbox{\mathbf{r}}_{t} represents the recurrent object reasoning state at time and \mbox{\mathbf{g}}_{t} is the global inter-frame interaction inferred at time such as described in Equation 4. In practice, this is implemented as a GRU, but for simplicity we omitted the gates in Equation (5). The pairwise mappings are implemented as an MLP. Figure 3 provides a visual explanation of the object head’s operating through time.
Our proposed ORN differs from [32] in three main points:
Objects have a semantic definition — we model relationships with respect to semantically meaningful entities (object instances) instead of feature map cells which do not have a semantically meaningful spatial extent. We will show in the experimental section that this is a key difference.
Objects are selected from different frames — we infer object pairwise relations only between objects present in two different sets. This is a key design choice which allows our model to reason about changes in object relationships over time.
Long range reasoning — integration of the object relations over time is recurrent by using a RNN for . Since reasoning from a full sequence cannot be done by inferring the relations between two frames, allows long range reasoning on sequences of variable length.
3.2 Object instance features
The object features \mbox{\mathbf{O}}_{t}=\big{\{}\mbox{\mathbf{o}}^{k}_{t}\big{\}}^{K}_{k=1} for each frame used for the ORN module described above are computed and collected from local regions predicted by a mask predictor. Independently for each frame \mbox{\mathbf{X}}_{t} of the input data block, we predict object instances as binary masks \mbox{\mathbf{B}}_{t}^{k} and associated object class predictions \mbox{\mathbf{c}}_{t}^{k}, a distribution over classes. We use Mask-RCNN [15], which is able to detect objects in a frame using region proposal networks [30] and produces a high quality segmentation mask for each object instance.
The objective is to collect features for each object instance, which jointly describe its appearance, the change in its appearance over time, and its shape, i.e. the shape of the binary mask. In theory, appearance could also be described by pooling the feature representation learned by the mask predictor (Mask R-CNN). However, in practice we choose to pool features from the dedicated object head such as shown in Figure 2, which also include motion through the spatio-temporal convolutions shared with the activity head:
[TABLE]
where \mbox{\mathbf{U}}_{t} is the feature map output by the object head, \mbox{\mathbf{u}}_{t}^{k} is a -dimensional vector of appearance and appearance change of object .
Shape information from the binary mask \mbox{\mathbf{B}}_{t}^{k} is extracted through the following mapping function: \mbox{\mathbf{b}}_{t}^{k}=g_{\phi}(\mbox{\mathbf{B}}_{t}^{k}), where is a MLP. Information about object in image \mbox{\mathbf{X}}_{t} is given by a concatenation of appearance, shape, and object class: \mbox{\mathbf{o}}_{t}^{k}=[\ \mbox{\mathbf{b}}_{t}^{k}\ \ \mbox{\mathbf{u}}_{t}^{k}\ \ \mbox{\mathbf{c}}_{t}^{k}\ ].
3.3 Global Motion and Context
Current approaches in video understanding focus on modeling the video from a high-level perspective. By a stack of spatio-temporal convolution and pooling they focus on learning global scene context information. Effective activity recognition requires integration of both of these sources: global information about the entire video content in addition to relational reasoning for making fine distinctions regarding object interactions and properties.
In our method, local low-level reasoning is provided through object head and the ORN module such as described above in Section 3.1. We complement this representation by high-level context information described by \mbox{\mathbf{V}}_{t} which are feature outputs from the activity head (orange block in Figure 2).
We use spatial global average pooling over \mbox{\mathbf{V}}_{t} to output -dimensional feature vectors denoted by \mbox{\mathbf{v}}_{t}, where \mbox{\mathbf{v}}_{t} corresponds to the context information of the video at timestep .
We model the dynamics of the context information through time by employing a RNN given by:
[TABLE]
where is the hidden state of and gives cues about the evolution of the context though time.
3.4 Recognition
Given an input video sequence \mbox{\mathbf{X}}_{1:T}, the two different streams corresponding to the activity head and the object head result in the two representations and , respectively where \mbox{\mathbf{h}}=\sum_{t}$$\mathbf{h}t and \mbox{\mathbf{r}}=\sum_{t}$$\mathbf{r}t. Each representation is the hidden state of the respective GRU, which were described in the preceding subsections. Recall that provides the global motion context while provides the object reasoning state output by the ORN module. We perform independent linear classification for each representation:
[TABLE]
where \mbox{\mathbf{y}}^{1},\mbox{\mathbf{y}}^{2} correspond to the logits from the activity head and the object head, respectively, and and are trainable weights (including biases). The final prediction is done by averaging logits \mbox{\mathbf{y}}^{1} and \mbox{\mathbf{y}}^{2} followed by softmax activation.
4 Network Architectures and feature dimensions
The input RGB images \mbox{\mathbf{X}}_{t} are of size where and correspond to the width and height and are of size 224 each. The object and activity heads (orange and green in Figure 2) are a joint convolutional neural network with Resnet50 architecture pre-trained on ImageNet/ILSVRC [31], with Conv1 and Conv5 blocks being inflated to 2.5D convolutions [44] (3D convolutions with a separable temporal dimension). This choice has been optimized on the validation set, as explained in Section 6 and shown in Table 5.
The last conv5 layers have been split into two different heads (activity head and object head). The intermediate feature representations \mbox{\mathbf{U}}_{t} and \mbox{\mathbf{V}}_{t} are of dimensions and , respectively. We provide a higher spatial resolution for the feature maps \mbox{\mathbf{U}}_{t} of the object head to get more precise local descriptors. This can be done by changing the stride of the initial conv5 layers from 2 to 1. Temporal convolutions have been configured to keep the same time temporal dimension through the network.
Global spatial pooling of activity features results in a 2048 dimensional feature vector fed into a GRU with 512 dimensional hidden state \mbox{\mathbf{s}}_{t}. ROI-Pooling of object features results in 2048 dimensional feature vectors \mbox{\mathbf{u}}_{t}^{k}. The encoder of the binary mask is a MLP with one hidden layer of size 100 and outputs a mask embedding \mbox{\mathbf{b}}_{t}^{k} of dimension 100. The number of object classes is 80, which leads in total to a 2229 dimensional object feature vector \mbox{\mathbf{o}}_{t}^{k}.
The non-linearity is implemented as an MLP with 2 hidden layers each with 512 units and produces an 512 dimensional output space. is implemented as a GRU with a 256 dimension hidden state \mbox{\mathbf{r}}_{t}. We use ReLU as the activation function after each layer for each network.
5 Training
We train the model with a loss split into two terms:
[TABLE]
where is the cross-entropy loss. The first term corresponds to supervised activity class losses comparing two different activity class predictions to the class ground truth: \hat{\mbox{\mathbf{y}}}^{1} is the prediction of the activity head, whereas \hat{\mbox{\mathbf{y}}}^{2} is the prediction of the object head, as given by Equations (8) and (9), respectively.
The second term is a loss which pushes the features of the object towards representations of the semantic object classes. The goal is to obtain features related to, both, motion (through the layers shared with the activity head), as well as as object classes. As ground-truth object classes are not available, we define the loss as the cross-entropy between the class label \mbox{\mathbf{c}}_{t}^{k} predicted by the mask predictor and a dedicated linear class prediction \hat{\mbox{\mathbf{c}}}_{t}^{k} based on features \mbox{\mathbf{u}}_{t}^{k}, which, as we recall, are RoI-pooled from :
[TABLE]
where trainable parameters (biases integrated) learned end-to-end together with the other parameters of the model.
We found that first training the object head only and then the full network was performing better. A ResNet50 network pretrained on ImageNet is modified by inflating some of its filters to 2.5 convolutions (3D convolutions with the time dimension separated), as described in Section 4; then by fine-tuning.
We train the model using the Adam optimizer [22] with an initial learning rate of on 30 epochs and use early-stopping criterion on the validation set for hyper-parameter optimization. Training takes 50 minutes per epoch on 4 Titan XP GPUs with clips of 8 frames.
6 Experimental results
We evaluated the method on three standard datasets, which represent difficult fine-grained activity recognition tasks: the Something-Something dataset, the VLOG dataset and the recently released EPIC Kitchens dataset.
**Something-Something (SS) **is a recent video classification dataset with 108,000 example videos and 157 classes [13]. It shows humans performing different actions with different objects, actions and objects being combined in different ways. Solving SS requires common sense reasoning and the state-of-the-art methods in activity recognition tend to fail, which makes this dataset challenging.
VLOG is a multi-label binary classification of human-object interactions recently released with 114,000 videos and 30 classes [12]. Classes correspond to objects, and labels of a class are 1 if a person has touched a certain object during the video, otherwise they are 0. It has recently been shown, that state-of-the-art video based methods [6] are outperformed on VLOG by image based methods like ResNet-50 [16], although these video methods outperform image based ResNet-50 on large-scale video datasets like the Kinetics dataset [6]. This suggests a gap between traditional datasets like Kinetics and the fine-grained dataset VLOG, making it particularly difficult.
**EPIC Kitchens (EPIC) **is an egocentric video dataset recently released containing 55 hours recording of daily activities [7]. This is the largest in first-person vision and the activities performed are non-scripted, which makes the dataset very challenging and close to real world data. The dataset is densely annotated and several tasks exist such as object detection, action recognition and action prediction. We focus on action recognition with 39’594 action segments in total and 125 actions classes (i.e verbs). Since the test set is not available yet we conducted our experiments on the training set (28’561 videos). We use the videos recorded by person 01 to person 25 for training (22’675 videos) and define the validation set as the remaining videos (5’886 videos).
For all datasets we rescale the input video resolution to . While training, we crop space-time blocks of spatial resolution and frames, with for the SS dataset and for VLOG and EPIC. We do not perform any other data augmentation. While training we extract frames from the entire video by splitting the video into sub-sequences and randomly sampling one frame per sub-sequence. The output sequence of size is called a clip. A clip aims to represent the full video with less frames. For testing we aggregate results of 10 clips. We use lintel [10] for decoding video on the fly.
The ablation study is done by using the train set as training data and we report the result on the validation set. We compare against other state-of-the-art approaches on the test set. For the ablation studies, we slightly decreased the computational complexity of the model: the base network (including activity and object heads) is a ResNet-18 instead of ResNet-50, a single clip of 4 frames is extracted from a video at test time.
**Comparison with other approaches. **Table 1 shows the performance of the proposed approach on the VLOG dataset. We outperform the state of the art on this challenging dataset by a margin of 4.2 points (44.7% accuracy against 40.5% by [16]). As mentioned above, traditional video approaches tend to fail on this challenging fine-grained dataset, providing inferior results. Table 4 shows performance on SS where we outperform the state of the art given by very recent methods (+2.3 points). On EPIC we re-implement standard baselines and report results on the validation set (Table 4) since the test set is not available. Our full method reports an accuracy of 40.89 and outperforms baselines by a large margin (+6.4 and +7.9 points respectively for against CNN-2D and I3D based on a ResNet-18).
**Effect of object-level reasoning. **Table 2 shows the importance of reasoning on the performance of the method. The baseline corresponds to the performance obtained by the activity head trained alone (inflated ResNet, in the ResNet-18 version for this table). No object level reasoning is present in this baseline. The proposed approach (third line) including an object head and the ORN module gains 0.8, 2.5 and 2.4 points compared to our baseline respectively on SS, on EPIC and on VLOG. This indicates that the reasoning module is able to extract complementary features compared to the activity head.
Using semantically defined objects proved to be important and led to a gain of 2 points on EPIC and 2.3 points on VLOG for the full model (6/12.7 points using the object head only) compared to an extension of Santoro et al. [32] operating on pixel level. This indicates importance of object level reasoning. The gain on SS is smaller (0.7 point with the full model and 7.8 points with the object head only) and can be explained by the difference in spatial resolution of the videos. Object detections and predictions of the binary masks are done using the initial video resolution. The mean video resolution for VLOG is and for EPIC is against for SS. Mask-RCNN has been trained on images of resolution and thus performs best on higher resolutions. The quality of the object detector is important for leveraging object level understanding then for the rest of the ablation study we focus on EPIC and VLOG datasets.
The function in Equation (5) is an important design choice in our model. In our proposed model, is recurrent over time to ensure that the ORN module captures long range reasoning over time, as shown in Equation (5). Removing the recurrence in this equation leads to an MLP instead of a (gated) RNN, as evaluated in row 4 of Table 2. Performance decreases by 1.1 point on VLOG and 1.4 points on EPIC. The larger gap for EPIC compared to VLOG and can arguably be explained by the fact that in SS actions cover the whole video, while solving VLOG requires detecting the right moment when the human-object interaction occurs and thus long range reasoning plays a less important role.
Visual features extracted from object regions are the most discriminative, however object shapes and labels also provide complementary information. Finally, the last part of Table 2 evaluates the effect of the cliques size for modeling the interactions between objects and show that pairwise cliques outperform cliques of size 1 and 3. We would like to recall, that even with unary cliques only, interactions between objects are still modeled. However, the model needs to find subspaces in the hidden representations associated to each interaction.
**CNN architecture and kernel inflations. **The convolutional architecture of the model was optimized over the validation set of the SS dataset, as shown in Table 5. The architecture itself (in terms of numbers of layers, filters etc.) is determined by pre-training on image classification. We optimized the choice of filter inflations from 2D to 2.5D or 3D for several convolutional blocks. This has been optimized for the single head model and using a ResNet-18 variant to speed up computation. Adding temporal convolutions increases performance up to 100% w.r.t. to pure 2D baselines. This indicates, without surprise, that motion is a strong cue. Inflating kernels to 2.5D on the input side and on the output side provided best performances, suggesting that temporal integration is required at a very low level (motion estimation) as well as on a very high level, close to reasoning. Our study also corroborates recent research in activity recognition, indicating that 2.5D kernels provide a good trade-off between high-capacity and learnable numbers of parameters. Finally temporal integration via RNN outperforms global average pooling over space and time. The choice of a (gated) RNN for temporal integration of the activity head features proved important (see Table 5) compared to global average pooling (GAP) over space and time.
**Visualizing the learned object interactions. **Figure 4 shows visualizations of the pairwise object relationships the model learned from data, in particular from the VLOG dataset. Each graph is computed for a given activity class, and strong arcs between two nodes in the graph indicate strong relationships between the object classes, i.e. the model detects a high correlation between these relationships and the activity. The graphs were obtained by thresholding the summed activations of each pairwise relationship in equation (4). Each pair can be assigned a pair of object classes \mbox{\mathbf{c}}_{t}^{j} and \mbox{\mathbf{c}}_{t}^{k} through the predictions of the object instance mask predictor. Integrating over all samples of the dataset for a given class leads to the visualizations in Figure 4. We can see that the object interactions are highly relevant to the detected activities: the person-touches-bed activity is correlated to interactions between relevant object classes person and bed. Similarly, activities human-bowl interaction and human-cup interaction show interactions with the respective objects bowl and cup. Moreover, other recovered relationships are highly correlated to the scene (for example, dining-table and bowl for activity human-bowl interaction).
Finally, Figure 5 shows some failure cases, which are either due to errors done by the object mask prediction (Mask R-CNN) or by the ORN itself.
7 Conclusion
We presented a method for activity recognition in videos which leverages object instance detections for visual reasoning on object interactions over time. The choice of reasoning over semantically well-defined objects is key to our approach and outperforms state of the art methods which reason on grid-levels, such as cells of convolutional feature maps. Temporal dependencies and causal relationships are dealt with by integrating relationships between different time instants. We evaluated the method on three difficult datasets, on which standard approaches do not perform well, and report state-of-the-art results.
Acknowledgements. This work was funded by grant Deepvision (ANR-15-CE23-0029, STPGP-479356-15), a joint French/Canadian call by ANR & NSERC.
The reference list from the paper itself. Each links out to its DOI / PubMed record.
- 1[1] Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., Vijayanarasimhan, S.: Youtube-8m: A large-scale video classification benchmark. ar Xiv preprint arxiv:1609.08675 (2016)
- 2[2] Baccouche, M., Mamalet, F., Wolf, C., Garcia, C., Baskurt, A.: Sequential deep learning for human action recognition. In: HBU (2011)
- 3[3] Baradel, F., Wolf, C., Mille, J., Taylor, G.: Glimpse clouds: Human activity recognition from unstructured feature points. In: CVPR (2018)
- 4[4] Battaglia, P.W., Pascanu, R., Lai, M., Rezende, D.J., Kavukcuoglu, K.: Interaction networks for learning about objects, relations and physics. In: NIPS (2016)
- 5[5] Bolei, Z., Zhang, A.A., Torralba, A.: Temporal relational reasoning in videos. In: ECCV (2018)
- 6[6] Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: CVPR (2017)
- 7[7] Damen, D., Doughty, H., Farinella, G.M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., Wray, M.: Scaling egocentric vision: The epic-kitchens dataset. In: ECCV (2018)
- 8[8] Deng, Z., Vahdat, A., Hu, H., Mori, G.: Structure inference machines: Recurrent neural networks for analyzing relations in group activity recognition. In: CVPR (2016)
