A Multimodal Dataset for Enhancing Industrial Task Monitoring and   Engagement Prediction

Naval Kishore Mehta; Arvind; Himanshu Kumar; Abeer Banerjee; Sumeet; Saurav; Sanjay Singh

arXiv:2501.05936·cs.CV·January 13, 2025

A Multimodal Dataset for Enhancing Industrial Task Monitoring and Engagement Prediction

Naval Kishore Mehta, Arvind, Himanshu Kumar, Abeer Banerjee, Sumeet, Saurav, Sanjay Singh

PDF

Open Access 1 Repo

TL;DR

This paper introduces the MIAM dataset, a comprehensive multimodal collection of industrial task videos with detailed annotations, and proposes a multimodal network to improve engagement prediction in human-robot collaboration.

Contribution

The paper presents a novel multimodal dataset capturing real-world industrial workflows and a fusion-based network for enhanced engagement prediction.

Findings

01

Improved accuracy in engagement state recognition

02

Multimodal data fusion enhances task monitoring

03

Dataset facilitates evaluation of action localization and object interaction

Abstract

Detecting and interpreting operator actions, engagement, and object interactions in dynamic industrial workflows remains a significant challenge in human-robot collaboration research, especially within complex, real-world environments. Traditional unimodal methods often fall short of capturing the intricacies of these unstructured industrial settings. To address this gap, we present a novel Multimodal Industrial Activity Monitoring (MIAM) dataset that captures realistic assembly and disassembly tasks, facilitating the evaluation of key meta-tasks such as action localization, object interaction, and engagement prediction. The dataset comprises multi-view RGB, depth, and Inertial Measurement Unit (IMU) data collected from 22 sessions, amounting to 290 minutes of untrimmed video, annotated in detail for task performance and operator behavior. Its distinctiveness lies in the integration of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

navalkishoremehta95/miam
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsOnline Learning and Analytics