HORIZON: A Benchmark for In-the-wild User Behaviour Modeling

Arnav Goel; Pranjal A Chitale; Bhawna Paliwal; Bishal Santra; Amit Sharma

arXiv:2604.17259·cs.IR·April 21, 2026

HORIZON: A Benchmark for In-the-wild User Behaviour Modeling

Arnav Goel, Pranjal A Chitale, Bhawna Paliwal, Bishal Santra, Amit Sharma

PDF

TL;DR

HORIZON is a comprehensive benchmark for user behavior modeling that emphasizes cross-domain, long-term, and generalizable scenarios, addressing limitations of existing narrow benchmarks.

Contribution

It introduces a large-scale, cross-domain dataset and new evaluation tasks that better reflect real-world user modeling challenges.

Findings

01

Current models struggle with cross-domain and long-term generalization.

02

Benchmarking reveals gaps between existing methods and real-world demands.

03

HORIZON provides a foundation for developing more robust user models.

Abstract

User behavior in the real world is diverse, cross-domain, and spans long time horizons. Existing user modeling benchmarks however remain narrow, focusing mainly on short sessions and next-item prediction within a single domain. Such limitations hinder progress toward robust and generalizable user models. We present HORIZON, a new benchmark that reformulates user modeling along three axes i.e. dataset, task, and evaluation. Built from a large-scale, cross-domain reformulation of Amazon Reviews, HORIZON covers 54M users and 35M items, enabling both pretraining and realistic evaluation of models in heterogeneous environments. Unlike prior benchmarks, it challenges models to generalize across domains, users, and time, moving beyond standard missing-positive prediction in the same domain. We propose new tasks and evaluation setups that better reflect real-world deployment scenarios. These…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.