SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

Kuangshi Ai; Haichao Miao; Kaiyuan Tang; Nathaniel Gorski; Jianxin Sun; Guoxi Liu; Helgi I. Ingolfsson; David Lenz; Hanqi Guo; Hongfeng Yu; Teja Leburu; Michael Molash; Bei Wang; Tom Peterka; Chaoli Wang; Shusen Liu

arXiv:2603.29139·cs.AI·April 1, 2026

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski, Jianxin Sun, Guoxi Liu, Helgi I. Ingolfsson, David Lenz, Hanqi Guo, Hongfeng Yu, Teja Leburu, Michael Molash, Bei Wang, Tom Peterka, Chaoli Wang, Shusen Liu

PDF

1 Repo 1 Datasets

TL;DR

SciVisAgentBench is a new comprehensive benchmark designed to evaluate scientific data analysis and visualization agents, enabling systematic comparison and progress tracking in realistic multi-step SciVis tasks.

Contribution

It introduces a structured taxonomy, a multimodal evaluation pipeline, and initial baselines, addressing the lack of principled benchmarks for SciVis agents.

Findings

01

Expert and LLM judges show good agreement in evaluations.

02

Baseline agents reveal significant capability gaps.

03

The benchmark supports systematic comparison and failure diagnosis.

Abstract

Recent advances in large language models (LLMs) have enabled agentic systems that translate natural language intent into executable scientific visualization (SciVis) tasks. Despite rapid progress, the community lacks a principled and reproducible benchmark for evaluating these emerging SciVis agents in realistic, multi-step analysis settings. We present SciVisAgentBench, a comprehensive and extensible benchmark for evaluating scientific data analysis and visualization agents. Our benchmark is grounded in a structured taxonomy spanning four dimensions: application domain, data type, complexity level, and visualization operation. It currently comprises 108 expert-crafted cases covering diverse SciVis scenarios. To enable reliable assessment, we introduce a multimodal outcome-centric evaluation pipeline that combines LLM-based judging with deterministic evaluators, including image-based…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

https://scivisagentbench.github.io
github

Datasets

SciVisAgentBench/SciVisAgentBench-tasks
dataset· 5.5k dl
5.5k dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.