Agent psychometrics: Task-level performance prediction in agentic coding benchmarks

Chris Ge; Daria Kryvosheieva; Daniel Fried; Uzay Girit; Kaivalya Hariharan

arXiv:2604.00594·cs.AI·April 2, 2026

Agent psychometrics: Task-level performance prediction in agentic coding benchmarks

Chris Ge, Daria Kryvosheieva, Daniel Fried, Uzay Girit, Kaivalya Hariharan

PDF

1 Repo

TL;DR

This paper introduces a framework that predicts individual task success in agentic coding benchmarks by augmenting Item Response Theory with rich task features and decomposing agent ability.

Contribution

It presents a novel method combining IRT with detailed task features and ability decomposition to predict task-level performance across diverse benchmarks.

Findings

01

Accurately predicts success on unseen tasks and benchmarks.

02

Enables better calibration of task difficulty for benchmark design.

03

Decomposes agent ability into LLM and scaffold components.

Abstract

As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environments, understanding which tasks will challenge agents and why becomes increasingly difficult. This is compounded by current practice: agent performance is typically measured by aggregate pass rates on benchmarks, but single-number metrics obscure the diversity of tasks within a benchmark. We present a framework for predicting success or failure on individual tasks tailored to the agentic coding regime. Our approach augments Item Response Theory (IRT) with rich features extracted from tasks, including issue statements, repository contexts, solutions, and test cases, and introduces a novel decomposition of agent ability into LLM and scaffold ability components. This parameterization enables us to aggregate evaluation data across heterogeneous…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

dariakryvosheieva/agent-psychometrics
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.