PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks

Yehoon Jang; Chaewon Lee; Hyun-seok Min; Sungchul Choi

arXiv:2601.04758·cs.CL·January 9, 2026

PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks

Yehoon Jang, Chaewon Lee, Hyun-seok Min, Sungchul Choi

PDF

Open Access 1 Video

TL;DR

PILOT-Bench is a new benchmark for evaluating legal reasoning in patent cases using IRAC-aligned classification tasks, revealing significant performance gaps between commercial and open-source language models.

Contribution

It introduces the first PTAB-centric benchmark with IRAC-aligned tasks for systematic evaluation of LLMs in patent legal reasoning.

Findings

01

Closed-source models outperform open-source models on Issue Type task.

02

Strong open-source model Qwen-8B achieves around 0.56 Micro-F1.

03

Benchmark highlights the need for improved reasoning in open-source LLMs.

Abstract

The Patent Trial and Appeal Board (PTAB) of the USPTO adjudicates thousands of ex parte appeals each year, requiring the integration of technical understanding and legal reasoning. While large language models (LLMs) are increasingly applied in patent and legal practice, their use has remained limited to lightweight tasks, with no established means of systematically evaluating their capacity for structured legal reasoning in the patent domain. In this work, we introduce PILOT-Bench, the first PTAB-centric benchmark that aligns PTAB decisions with USPTO patent data at the case-level and formalizes three IRAC-aligned classification tasks: Issue Type, Board Authorities, and Subdecision. We evaluate a diverse set of closed-source (commercial) and open-source LLMs and conduct analyses across multiple perspectives, including input-variation settings, model families, and error tendencies.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks· underline

Taxonomy

TopicsIntellectual Property and Patents · Explainable Artificial Intelligence (XAI) · Law, AI, and Intellectual Property