ElecBench: a Power Dispatch Evaluation Benchmark for Large Language   Models

Xiyuan Zhou; Huan Zhao; Yuheng Cheng; Yuji Cao; Gaoqi Liang; Guolong; Liu; Wenxuan Liu; Yan Xu; Junhua Zhao

arXiv:2407.05365·cs.AI·August 13, 2024·3 cites

ElecBench: a Power Dispatch Evaluation Benchmark for Large Language Models

Xiyuan Zhou, Huan Zhao, Yuheng Cheng, Yuji Cao, Gaoqi Liang, Guolong, Liu, Wenxuan Liu, Yan Xu, Junhua Zhao

PDF

Open Access 1 Repo

TL;DR

ElecBench is a comprehensive evaluation benchmark designed to assess large language models' performance in the power sector across various professional and general scenarios, aiming to facilitate technological progress and application.

Contribution

This paper introduces ElecBench, the first specialized benchmark for evaluating LLMs in the power sector, covering sector-specific scenarios and multiple performance metrics.

Findings

01

Evaluated 8 LLMs across diverse scenarios and metrics.

02

Provided a public test set for transparent benchmarking.

03

Identified strengths and limitations of current LLMs in power applications.

Abstract

In response to the urgent demand for grid stability and the complex challenges posed by renewable energy integration and electricity market dynamics, the power sector increasingly seeks innovative technological solutions. In this context, large language models (LLMs) have become a key technology to improve efficiency and promote intelligent progress in the power sector with their excellent natural language processing, logical reasoning, and generalization capabilities. Despite their potential, the absence of a performance evaluation benchmark for LLM in the power sector has limited the effective application of these technologies. Addressing this gap, our study introduces "ElecBench", an evaluation benchmark of LLMs within the power sector. ElecBench aims to overcome the shortcomings of existing evaluation benchmarks by providing comprehensive coverage of sector-specific scenarios,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

xiyuan-zhou/elecbench-a-power-dispatch-evaluation-benchmark-for-large-language-models
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling

MethodsSparse Evolutionary Training