CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat, Divyansh Agarwal, Prafulla Kumar Choubey, Yixin Mao, Silvio Savarese, Caiming Xiong, Chien-Sheng Wu

TL;DR
CRMArena-Pro is a comprehensive benchmark designed to evaluate LLM agents across diverse real-world business scenarios, revealing significant gaps in current AI capabilities for enterprise applications.
Contribution
The paper introduces CRMArena-Pro, a new benchmark with expert-validated tasks, multi-turn interactions, and confidentiality assessments, addressing limitations of existing benchmarks in business AI evaluation.
Findings
Top LLM agents achieve only 58% success in single-turn tasks
Performance drops to around 35% in multi-turn interactions
Agents show near-zero inherent confidentiality awareness
Abstract
While AI agents hold transformative potential in business, effective performance benchmarking is hindered by the scarcity of public, realistic business data on widely used platforms. Existing benchmarks often lack fidelity in their environments, data, and agent-user interactions, with limited coverage of diverse business scenarios and industries. To address these gaps, we introduce CRMArena-Pro, a novel benchmark for holistic, realistic assessment of LLM agents in diverse professional settings. CRMArena-Pro expands on CRMArena with nineteen expert-validated tasks across sales, service, and 'configure, price, and quote' processes, for both Business-to-Business and Business-to-Customer scenarios. It distinctively incorporates multi-turn interactions guided by diverse personas and robust confidentiality awareness assessments. Experiments reveal leading LLM agents achieve only around 58%…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsBusiness Process Modeling and Analysis · Big Data and Business Intelligence · Open Source Software Innovations
Methodstravel james
