ObjectVisA-120: Object-based Visual Attention Prediction in Interactive Street-crossing Environments

Igor Vozniak; Philipp Mueller; Nils Lipp; Janis Sprenger; Konstantin Poddubnyy; Davit Hovhannisyan; Christian Mueller; Andreas Bulling; Philipp Slusallek

arXiv:2601.13218·cs.CV·February 2, 2026

ObjectVisA-120: Object-based Visual Attention Prediction in Interactive Street-crossing Environments

Igor Vozniak, Philipp Mueller, Nils Lipp, Janis Sprenger, Konstantin Poddubnyy, Davit Hovhannisyan, Christian Mueller, Andreas Bulling, Philipp Slusallek

PDF

Open Access

TL;DR

This paper introduces ObjectVisA-120, a comprehensive VR dataset for object-based visual attention in street-crossing scenarios, along with a new evaluation metric and a novel graph-based attention prediction model, advancing understanding and modeling of human attention.

Contribution

The paper provides a new VR dataset with rich annotations for object-based attention, a novel metric oSIM for evaluation, and a graph-encoded attention model, SUMGraph, improving attention prediction performance.

Findings

01

Optimizing for object-based attention improves oSIM scores.

02

SUMGraph outperforms state-of-the-art models in attention prediction.

03

ObjectVisA-120 enables better evaluation of object-based attention models.

Abstract

The object-based nature of human visual attention is well-known in cognitive science, but has only played a minor role in computational visual attention models so far. This is mainly due to a lack of suitable datasets and evaluation metrics for object-based attention. To address these limitations, we present ObjectVisA-120 -- a novel 120-participant dataset of spatial street-crossing navigation in virtual reality specifically geared to object-based attention evaluations. The uniqueness of the presented dataset lies in the ethical and safety affiliated challenges that make collecting comparable data in real-world environments highly difficult. ObjectVisA-120 not only features accurate gaze data and a complete state-space representation of objects in the virtual environment, but it also offers variable scenario complexities and rich annotations, including panoptic segmentation, depth…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsVisual Attention and Saliency Detection · Gaze Tracking and Assistive Technology · Multimodal Machine Learning Applications