FormulaCode: Evaluating Agentic Optimization on Large Codebases

Atharva Sehgal; James Hou; Akanksha Sarkar; Ishaan Mantripragada; Swarat Chaudhuri; Jennifer J. Sun; Yisong Yue

arXiv:2603.16011·cs.SE·May 18, 2026

FormulaCode: Evaluating Agentic Optimization on Large Codebases

Atharva Sehgal, James Hou, Akanksha Sarkar, Ishaan Mantripragada, Swarat Chaudhuri, Jennifer J. Sun, Yisong Yue

PDF

2 Repos 1 Datasets

TL;DR

FormulaCode introduces a comprehensive benchmark with real-world Python codebases and multi-objective metrics to evaluate large language model agents' ability to optimize entire repositories.

Contribution

It provides the first large-scale, multi-objective benchmark for assessing agentic optimization on real-world code repositories, addressing limitations of synthetic task-based benchmarks.

Findings

01

LLM agents struggle with holistic, multi-objective codebase optimization.

02

Existing benchmarks do not adequately evaluate real-world optimization capabilities.

03

FormulaCode enables more realistic assessment of LLM agent performance.

Abstract

Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to optimize entire codebases under realistic constraints. Existing code benchmarks largely rely on synthetic tasks, binary correctness signals, or single-objective evaluation, limiting their ability to assess holistic optimization behavior. We introduce FormulaCode, a benchmark for evaluating agentic optimization on large, real-world codebases with fine-grained, multi-objective performance metrics. FormulaCode comprises 957 performance bottlenecks mined from scientific Python repositories on GitHub, each paired with expert-authored patches and, on average, 264.6 community-maintained performance workloads per task, enabling the holistic ability of LLM agents to optimize codebases under realistic correctness and performance constraints. Our evaluations…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Datasets

formulacode/formulacode-all
dataset· 120 dl
120 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMachine Learning in Materials Science · Natural Language Processing Techniques · Topic Modeling