Can LLM Already Serve as A Database Interface? A BIg Bench for   Large-Scale Database Grounded Text-to-SQLs

Jinyang Li; Binyuan Hui; Ge Qu; Jiaxi Yang; Binhua Li; Bowen Li,; Bailin Wang; Bowen Qin; Rongyu Cao; Ruiying Geng; Nan Huo; Xuanhe Zhou,; Chenhao Ma; Guoliang Li; Kevin C.C. Chang; Fei Huang; Reynold Cheng; Yongbin; Li

arXiv:2305.03111·cs.CL·November 16, 2023·82 cites

Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs

Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li,, Bailin Wang, Bowen Qin, Rongyu Cao, Ruiying Geng, Nan Huo, Xuanhe Zhou,, Chenhao Ma, Guoliang Li, Kevin C.C. Chang, Fei Huang, Reynold Cheng, Yongbin, Li

PDF

Open Access 1 Repo 10 Models 3 Datasets 1 Video

TL;DR

This paper introduces Bird, a large-scale benchmark for text-to-SQL tasks on extensive databases, highlighting challenges like database value comprehension and demonstrating that current models like ChatGPT still have significant room for improvement in real-world scenarios.

Contribution

The paper presents Bird, a comprehensive benchmark with large-scale databases, emphasizing database value understanding and providing insights into model performance and efficiency in real-world applications.

Findings

01

ChatGPT achieves only 40.08% execution accuracy on Bird.

02

Database values significantly impact text-to-SQL accuracy.

03

Challenges remain in scaling text-to-SQL models for large databases.

Abstract

Text-to-SQL parsing, which aims at converting natural language instructions into executable SQLs, has gained increasing attention in recent years. In particular, Codex and ChatGPT have shown impressive results in this task. However, most of the prevalent benchmarks, i.e., Spider, and WikiSQL, focus on database schema with few rows of database contents leaving the gap between academic study and real-world applications. To mitigate this gap, we present Bird, a big benchmark for large-scale database grounded in text-to-SQL tasks, containing 12,751 pairs of text-to-SQL data and 95 databases with a total size of 33.4 GB, spanning 37 professional domains. Our emphasis on database values highlights the new challenges of dirty database contents, external knowledge between NL questions and database contents, and SQL efficiency, particularly in the context of massive databases. To solve these…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

bird-bench/mini_dev
none

Models

Datasets

Videos

Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs· slideslive

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Software Engineering Research