GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning

Zhiheng Jiang; Yunzhe Wang; Ryan Marr; Ellen Novoseller; Benjamin T. Files; Volkan Ustun

arXiv:2601.20753·cs.LG·February 4, 2026

GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning

Zhiheng Jiang, Yunzhe Wang, Ryan Marr, Ellen Novoseller, Benjamin T. Files, Volkan Ustun

PDF

Open Access

TL;DR

GraphAllocBench introduces a versatile, graph-based benchmark for preference-conditioned multi-objective policy learning, enabling realistic, scalable evaluation of algorithms in complex resource allocation scenarios with new metrics for preference consistency.

Contribution

It presents a novel graph-based resource allocation benchmark with diverse objectives and preferences, along with new evaluation metrics, facilitating advanced research in preference-conditioned reinforcement learning.

Findings

01

Existing MORL approaches have limitations exposed by GraphAllocBench.

02

Graph Neural Networks show promise in complex, high-dimensional allocation tasks.

03

The benchmark's flexibility allows for comprehensive evaluation of preference-conditioned policies.

Abstract

Preference-Conditioned Policy Learning (PCPL) in Multi-Objective Reinforcement Learning (MORL) aims to approximate diverse Pareto-optimal solutions by conditioning policies on user-specified preferences over objectives. This enables a single model to flexibly adapt to arbitrary trade-offs at run-time by producing a policy on or near the Pareto front. However, existing benchmarks for PCPL are largely restricted to toy tasks and fixed environments, limiting their realism and scalability. To address this gap, we introduce GraphAllocBench, a flexible benchmark built on a novel graph-based resource allocation sandbox environment inspired by city management, which we call CityPlannerEnv. GraphAllocBench provides a rich suite of problems with diverse objective functions, varying preference conditions, and high-dimensional scalability. We also propose two new evaluation metrics -- Proportion of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Advanced Multi-Objective Optimization Algorithms · Explainable Artificial Intelligence (XAI)