Analysis of the Reasoning with Redundant Information Provided Ability of   Large Language Models

Wenbei Xie

arXiv:2310.04039·cs.CL·October 9, 2023·1 cites

Analysis of the Reasoning with Redundant Information Provided Ability of Large Language Models

Wenbei Xie

PDF

Open Access

TL;DR

This paper introduces the RRIP benchmark to evaluate LLM reasoning with redundant information, revealing current models struggle with such tasks and highlighting the need for training data improvements.

Contribution

The study proposes a new RRIP benchmark and modified GSM-8K dataset to assess LLM reasoning with redundant info, exposing limitations of current models.

Findings

01

Models perform poorly on RRIP tasks compared to standard benchmarks.

02

Performance declines are significant when handling redundant information.

03

Training models with redundant data could improve reasoning abilities.

Abstract

Recent advancements in Large Language Models (LLMs) have demonstrated impressive capabilities across a range of natural language processing tasks, especially in reasoning, a cornerstone for achieving Artificial General Intelligence (AGI). However, commonly used benchmarks may not fully encapsulate the inferential abilities of these models in real-world scenarios. To address this gap, a new form of Question-Answering (QA) task, termed Reasoning with Redundant Information Provided (RRIP), is introduced. The study designed a modified version of the grade school math 8K (GSM-8K) dataset which has several variants focusing on different attributes of redundant information. This investigation evaluates two popular LLMs, LlaMA2-13B-chat and generative pre-trained transformer 3.5 (GPT-3.5), contrasting their performance on traditional QA tasks against the RRIP tasks. Findings indicate that while…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Text Readability and Simplification

MethodsFocus